Use the cost-saving playbook to evaluate these dated candidates on equal tasks; token prices alone omit failures and review time.
Budget models are useful when their lower token price survives retries, long outputs, and human review. The relevant metric is cost per successful task, not the cheapest input row.
September 27 shortlist
Start with two candidates whose access route fits your work. This is an editorial shortlist, not a measured ranking:
| Candidate | Why evaluate it | Read before paying |
|---|---|---|
| MiMo V2.6 Flash | Low listed API prices, multimodal input and released weights | API, Token Plan and Desktop are separate products; self-hosting has substantial hardware costs |
| DeepSeek V4.1 Flash | Direct API value and native vision | Use the exact current model ID and pricing window; 0731 results are historical |
| GLM-5.3 | Another coding and agent candidate | Check Z.AI’s current access and plan rules; GLM-5.2 scores below do not evaluate 5.3 |
| GPT-6 Luna | Current OpenAI bounded-work and extraction route | Separate included subscription use from API billing and count review effort; verify the live owner price |
Prices and benchmark details stay on their linked owners. See how to compare accepted-result costs. AIHackers has no MiMo V2.6 CAR result.
Current routing
- Start with GPT-6 Luna when the OpenAI route fits; the current provider price anchor is $0.10/$0.50 input/output with $0.01 cache read.
- Use GPT-6 Sol or Opus 5.5 when retries, review, or task risk justify a premium route.
- Read the GPT-6 Astra planning evaluation before using Astra outside planning tasks.
August 1 snapshot: historical comparison
Everything below preserves the August 1 review, including its model names, prices, rankings and recommendations. It is not a current price list. In particular, MiMo V2.5 and DeepSeek 0731 have newer revisions in the shortlist above.
Historical shortlist
| Model | Input / output per 1M | Context | Current role |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | $0.14 cache miss, $0.0028 cache hit / $0.28 | 1M | Direct-API and MIT-weight value candidate; verify accepted-task behavior |
| Kimi K2.7 Code | $0.95 cache miss, $0.19 cache hit / $4.00 | 256K | Cheaper routine Kimi coding API lane |
| Gemini 3 Flash | $0.50 / $3.00 | ~1M | Gemini high-context value lane |
| MiniMax M3 | Standard PAYG starts at $0.60 / $2.40 | 1M | Coding-agent value lane to test |
| Xiaomi MiMo-V2.5 | Overseas list starts at $0.14 / $0.28 | 1M | Low-cost long-context and open-weight evaluation lane |
| GPT-5.6 Luna | $0.20 / $1.20 | 1.05M | Generally available high-volume GPT-5.6 API tier; long-context rates apply above 272K input |
Subscription and Token Plan quotas are not API prices. Compare them separately and verify checkout before purchase.
Historical August starting points
Kimi K2.7 Code
Use K2.7 Code for lower-cost Kimi coding searches. Moonshot documents base and HighSpeed IDs, 256K context, required thinking mode, multimodal input, and automatic context caching. Use Kimi K3 for newest-Kimi, 1M-context, or K3 benchmark intent.
Kimi’s published improvements over K2.6 are vendor evidence. AIHackers has not verified a normalized independent K2.7 coding score or run a controlled repository comparison.
Gemini 3 Flash
Use Gemini when its API or Vertex AI route, context size, and tool support fit. Verify the current model stage, regional availability, and exact free/paid rates; preview and free-tier limits can change faster than this page.
MiniMax M3 and Xiaomi MiMo
Both are newer low-cost, long-context lanes worth testing. Their strongest coding claims are vendor-reported, so require the same real repository tasks and pass conditions used for Kimi, GLM, or Claude.
GPT-5.6 Luna: historical route
The August review treated GPT-5.6 Luna as the OpenAI value default for bounded Codex work and high-volume API tasks. Its message ranges and July price change remain historical evidence. The historical value-default analysis keeps subscription allowance and API billing separate.
DeepSeek V4 Flash 0731
DeepSeek’s direct API maps deepseek-v4-flash to the official 0731 release. Artificial Analysis scores it 50, one point behind Luna max, and places it on the evaluator’s Intelligence-versus-Cost frontier. The 0731 analysis keeps that independent finding separate from vendor benchmarks, the older preview’s 0/10 security result, and AIHackers’ site-owned not-run state.
August evidence table
benchmark artifact
Budget and Value Model Evidence
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | DeepSeek |
historical Historical 0731 scores and prices. The direct legacy Flash name now serves V4.1 Flash; 0731 MIT weights remain a separate artifact. | 1M | $0.14 / 1M Historical July 31 pricing: $0.14 cache-miss input / $0.0028 cache-hit input / $0.28 output per 1M tokens | $0.28 / 1M | Artificial Analysis reports Intelligence Index 50, one point behind GPT-5.6 Luna at max effort; this is shortlist evidence, not a universal coding-quality claim. | DeepSeek documents tool calls, Responses API, and Anthropic-compatible access; real harness behavior remains workload-specific. |
| Independent write-up reports about 206M total output tokens for the Intelligence Index; production speed and accepted-task efficiency are not verified. | Historical 0731 evidence and separate MIT weights; use the V4.1 Flash owner guide for current direct-API routing. Neither the preview test nor 0731 benchmarks evaluate V4.1. | DeepSeek V4 Flash 0731 update [archive], DeepSeek V4 Flash 0731 model card [archive], DeepSeek API pricing [archive], Artificial Analysis DeepSeek V4 Flash 0731, OpenCode Zen pricing [archive], Kasra Rahjerdi’s original June 3 field test, AIHackers analysis of Kasra Rahjerdi’s field test | 2026-08-01 |
| GPT-5.6 Luna | OpenAI |
historical Previous generation; dated scores and prices retained. Current routing: GPT-6 Astra, Sol and Luna. Historical status here does not imply API retirement. | 1.05M | $0.20 / 1M $0.20 input / $0.02 cache read / $1.20 output per 1M tokens | $1.20 / 1M | Artificial Analysis reports 75 on its Coding Agent Index at max effort; BenchLM ranks Luna #6/129 in coding with an Estimated overall position. | Generally available in API and paid Codex plans. |
| Vendor-positioned as fastest; measured production latency is not verified. | Historical comparison record; use GPT-6 Astra, Sol and Luna for current evaluation candidates. | OpenAI GPT-5.6 general availability [archive], OpenAI GPT-5.6 price update [archive], OpenAI API pricing [archive], OpenAI GPT-5.6 Luna model page [archive], Artificial Analysis GPT-5.6 evaluation [archive], Agent Arena leaderboard, BenchLM GPT-5.6 Luna profile [archive], OpenAI GPT-5.6 availability [archive], OpenAI GPT-5.6 system card | 2026-08-01 |
| Kimi K2.7 Code | Moonshot AI |
active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M $0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| Gemini 3 Flash |
active Current Gemini value lane where Gemini API or Vertex AI fits. | 1.05M input | $0.50 / 1M $0.50 input / $3.00 output per 1M tokens | $3.00 / 1M | not verified | Function calling and code execution supported. | not verified | Preview model positioned for lower latency; independent value not imported. | High-context value lane when Gemini API or Vertex AI fits. | Gemini API models, Gemini API pricing, Artificial Analysis: Gemini 3 Flash, LMArena leaderboard dataset | 2026-05-26 | |
| GLM-5.2 | Z.AI |
historical Historical July comparison; Z.AI now documents GLM-5.3. These scores and prices do not evaluate 5.3. Historical status does not imply API retirement. | 1M | $1.40 / 1M $1.40 input / $0.26 cached input / $4.40 output per 1M tokens | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | Dated July value candidate; retain for pinned integrations and historical comparisons. Check the newer GLM revision separately before new routing decisions. | Z.AI GLM-5.3 successor documentation, Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
DeepSeek 0731 and GPT-5.6 Luna are shortlist candidates; local quality, latency, and cost-per-successful-task results remain not-run or not verified.
GLM-5.2 sits above the strict $1 input threshold at $1.40, but it is the relevant value comparison because it adds 1M context and independent Artificial Analysis evidence.
What Not to Compare Directly
- SWE-bench Verified and SWE-Bench Pro.
- A vendor’s internal preference test and an independent aggregate index.
- Base API price and a monthly coding-plan quota.
- Cache-hit price and uncached first-run price.
- A model’s context limit and its ability to use that context accurately.
Do not claim one model delivers a percentage of another model’s total capability from a single benchmark.
Repository Test
Run each accessible model on:
| Task | Pass condition |
|---|---|
| Repository map | Correct modules and commands; no invented files |
| Bug fix | Minimal patch and relevant passing tests |
| Refactor | Preserved behavior and repository conventions |
| Review | Concrete file-grounded findings |
Record exact model ID, settings, prompt, tokens, cache behavior, retries, latency, accepted patch, and review time. Models that are inaccessible or provider-managed should be labeled accordingly rather than assigned estimated results.
Historical Budget Context
Kimi K2.5 remains useful for older free-hosting and integration searches. Kimi K2.6 remains relevant for compatible non-thinking and multimodal workflows. Neither should replace K3 for newest-Kimi intent or K2.7 Code for cheaper Kimi coding.
Older access pages can retain K2.5 when that exact hosted model is still offered. They must not describe it as Kimi’s latest coding model.
Historical August verdict
- Start with Kimi K2.7 Code when Kimi’s API and 256K context fit.
- Start with DeepSeek V4 Flash 0731 when direct-API price or MIT weights are the main constraint.
- Test GLM-5.2 when 1M context and independent shortlist evidence justify slightly higher input pricing.
- Evaluate Gemini 3 Flash, MiniMax M3, or Xiaomi MiMo when their provider/tool route fits.
- Test GPT-5.6 Luna when the OpenAI API or paid Codex route fits the workload; this is the historical August recommendation.
- Escalate to Opus 4.8 only when a premium second pass materially improves the outcome; this is the historical August recommendation.
Sources
- Kimi: K2.7 Code quickstart and pricing
- Z.AI: GLM-5.2 and pricing
- Google: Gemini models and pricing
- OpenAI: GPT-6 Sol/Luna pricing owner and GPT-6 Astra planning evidence
- OpenAI: historical GPT-5.6 general availability
- OpenAI: historical July 30 price update and API pricing
- Artificial Analysis: GLM-5.2
- DeepSeek: 0731 model card and direct pricing
- Artificial Analysis: DeepSeek V4 Flash 0731
Related links
- /models/kimi-k2.7-code/
- /models/kimi-k3/
- /models/glm-5.2/
- /value/gpt-6-sol-luna-pricing/
- /posts/gpt-6-astra-planning-value/
- /models/gpt-5-6/ — historical GPT-5.6 family guide
- /posts/gpt-5-6-luna-price-cut/ — historical Luna price-cut analysis
- /posts/deepseek-v4-flash-0731-value-frontier/
- /compare/models/mid-range/
- /value/free-stack/
Historical comparison verified August 1, 2026; GPT-6 and Opus 5.5 pointers in the linked shortlist were refreshed September 27. Other provider entries retain their source dates; prices, access and benchmark versions can change independently.