October 2 routing correction: This page preserves its July evidence and recommendations. Z.AI now documents GLM-5.3; GLM-5.2 scores and prices below do not evaluate that revision. Historical status here does not imply API retirement. Use the selected model hub before choosing a new route.
Short verdict: Start with GLM-5.2 when the job needs 1M context, whole-repo context retention, or Z.AI-supported coding-tool routing at a lower token price. Start with Kimi K2.7 Code when you want cheaper Kimi coding economics, multimodal tool calls, or a HighSpeed API lane. Use Kimi K3 when the question is newest-Kimi or frontier-adjacent Kimi performance, and keep Kimi K2.6 in the comparison because it still has exact-match demand and remains relevant where non-thinking mode or existing Kimi integrations matter.
This is a shortlist guide, not a leaderboard. Vendor benchmark claims can help pick what to test, but the winning model is the one that fixes your actual bug with fewer retries and less human cleanup.
Quick Facts
| Model | Current role | Context | Input price | Output price |
|---|---|---|---|---|
| GLM-5.2 | Z.AI July flagship comparison; now historical | 1M | $1.40 / 1M, $0.26 cached | $4.40 / 1M |
| Kimi K3 | Kimi newest flagship | 1M | $0.30 cache hit / $3.00 cache miss | $15.00 / 1M |
| Kimi K2.7 Code | Cheaper Kimi coding release | 256K-class | $0.19 cache hit / $0.95 cache miss | $4.00 / 1M |
| Kimi K2.7 Code HighSpeed | Faster K2.7 API lane | 256K-class | $0.38 cache hit / $1.90 cache miss | $8.00 / 1M |
| Kimi K2.6 | Prior Kimi multimodal/API comparison lane | 256K-class | $0.16 cache hit / $0.95 cache miss | $4.00 / 1M |
Which Should You Test First?
| If your constraint is… | Start with | Why |
|---|---|---|
| Whole-repo context or very long tasks | GLM-5.2 | Z.AI documents 1M context and 128K max output |
| Lowest cache-hit input cost | Kimi K2.6 or Kimi K2.7 Code | Kimi cache-hit input is lower than GLM’s public non-cached input anchor |
| Cheaper dedicated Kimi coding API | Kimi K2.7 Code | K2.7 remains the lower-cost 256K coding lane; K3 is the newest flagship |
| Fast coding responses | Kimi K2.7 Code HighSpeed | Same K2.7 model, higher listed output speed, higher token prices |
| Supported-tool subscription economics | GLM Coding Plan | Z.AI sells a GLM Coding Plan for supported coding tools |
| Non-thinking Kimi mode | Kimi K2.6 | K2.7 Code requires thinking; K2.6 can disable it |
| Multimodal coding inputs | Kimi K2.7 Code or K2.6 | Kimi docs emphasize text, image, and video input |
Price And Context
GLM-5.2’s price case is not “cheapest raw token.” Its case is 1M context plus supported coding-tool economics. Kimi’s price case is cheaper input in cache-hit workflows and a simple OpenAI-compatible API path, with K2.7 Code HighSpeed available when speed is worth the higher output price.
| Buyer question | GLM-5.2 answer | Kimi answer |
|---|---|---|
| “Which has bigger context?” | GLM-5.2 at 1M | Kimi K2.7/K2.6 at 256K-class |
| “Which has cheaper output?” | $4.40 / 1M | $4.00 / 1M base K2.7/K2.6; $8.00 HighSpeed |
| “Which has cheaper cache-hit input?” | $0.26 / 1M cached input | $0.19 K2.7 Code, $0.16 K2.6 |
| “Which has a subscription lane?” | GLM Coding Plan | Kimi membership/Kimi Code paths, but verify current quota wording |
| “Which should handle routine work?” | Test GLM if tool support fits | Test Kimi if API/tooling support fits |
Tool Fit
| Workflow | Better first test | Reason |
|---|---|---|
| Claude Code-style supported-tool routing | GLM-5.2 via Z.AI | Z.AI documents Anthropic-compatible and supported-tool routes |
| OpenAI-compatible API experiments | Both | Both have OpenAI-compatible API paths |
| Kimi-native latest flagship | Kimi K3 / Kimi Code | Kimi-specific workflows should test K3 separately from cheaper K2.7 routing |
| Kimi-native budget coding | Kimi K2.7 Code / Kimi Code | Kimi-specific workflows can keep K2.7 when price matters more than newest-model testing |
| OpenClaw/BYOK agent routing | Both | Use provider-specific docs and treat subscriptions separately from direct API billing |
| Very large repo audit | GLM-5.2 | 1M context is the main spec advantage |
| Vision/video coding task | Kimi K2.7 Code or K2.6 | Kimi docs explicitly emphasize multimodal input |
Caveats
- Z.AI benchmark language is vendor-published. Use it to decide what to test, not to claim “best model.”
- Kimi K2.7 Code benchmark and speed language is also vendor-published. Measure your own cost per accepted patch.
- Kimi K2.7 Code requires thinking mode. If your workflow needs non-thinking behavior, test K2.6.
- Subscription quota and membership wording can move faster than API pricing docs. Treat checkout and current docs as final before purchase.
- Context size is not quality. A smaller-context model can still win if it follows your repo conventions better.
Benchmark Evidence
benchmark artifact
GLM-5.2 vs Kimi Evidence
| Model | Provider | Status | Context | Input price | Output price | Coding signal | Tool-use signal | Benchmark evidence | Speed | Verdict | Sources | Checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GLM-5.2 | Z.AI |
historical Historical July comparison; Z.AI now documents GLM-5.3. These scores and prices do not evaluate 5.3. Historical status does not imply API retirement. | 1M | $1.40 / 1M $1.40 input / $0.26 cached input / $4.40 output per 1M tokens | $4.40 / 1M | Z.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1. | Supported-tool coding lane; BFCL score not imported. |
| Artificial Analysis flags higher output-token use; measure total cost per successful task. | Dated July value candidate; retain for pinned integrations and historical comparisons. Check the newer GLM revision separately before new routing decisions. | Z.AI GLM-5.3 successor documentation, Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| Kimi K3 | Moonshot AI |
active Kimi API/product flagship; Kimi K3 License weights. NVIDIA also lists K3 under trial terms as checked September 14; account access untested. | 1M | $3.00 / 1M $0.30 cache-hit / $3.00 cache-miss input / $15.00 output per 1M tokens | $15.00 / 1M | Moonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified. | Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching. |
| Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task. | Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests. | NVIDIA Kimi K3 model card, Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive] | 2026-08-01 |
| Kimi K2.7 Code | Moonshot AI |
active Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices. | 256K | $0.95 / 1M $0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates | $4.00 / 1M | Kimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported. | OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart. |
| HighSpeed model ID exists at a higher token price; latency not independently measured here. | Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough. | Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| Kimi K2.6 | Moonshot AI |
historical Prior Kimi comparison lane retained for compatible integrations. | 256K | $0.95 / 1M $0.16 cache-hit input / $0.95 cache-miss input / $4.00 output per 1M tokens | $4.00 / 1M | not verified | ToolCalls and OpenAI-compatible API supported; BFCL score not imported. | not verified | not verified | Prior Kimi comparison baseline where K2.6-specific integrations or multimodal behavior matter. | Kimi models, Kimi K2.6 quickstart, Kimi K2.6 pricing, Artificial Analysis: Kimi K2.6, SWE-bench, Berkeley Function Calling Leaderboard | 2026-06-28 |
| Claude Opus 4.8 | Anthropic |
historical Historical comparison; use Opus 5.5 for current premium evaluation. | 1M | $5.00 / 1M $5.00 input / $25.00 output per 1M tokens | $25.00 / 1M | Historical premium Claude baseline; use Opus 5 for new task-level comparisons. | Still available for pinned integrations; new Claude premium routing should test Opus 5. |
| Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary. | Historical premium baseline. Use Claude Opus 5.5 for current Claude premium routing. | Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard | 2026-07-25 |
Kimi K3 has early independent evidence and vendor launch evidence. GLM-5.2 has an independent Artificial Analysis aggregate plus vendor coding scores. K2.7 results are vendor-reported relative deltas versus K2.6. Do not normalize them into one leaderboard.
Eval Checklist
Run the same tasks through each candidate:
| Eval | What to measure | Winner signal |
|---|---|---|
| Bug fix | One real failing test | Smallest correct patch, fewest retries |
| Refactor | 2-4 files with local style constraints | Preserves behavior and tests |
| Repo audit | Architecture map from docs and code | Accurate module boundaries |
| Tool calls | Multi-step task with tool results | Keeps reasoning/tool context intact |
| Cost | Actual cache-hit/cache-miss and output usage | Lower cost per accepted result |
| Review | Patch review against a known risky change | Concrete findings without invented policy |
Related links
- /models/glm-5.2/ - GLM-5.2 model guide
- /models/kimi-k2.7-code/ - Kimi K2.7 Code model guide
- /models/kimi-k3/ - Kimi K3 model guide
- /tools/zai/ - Z.AI Coding Plan setup and caveats
- /tools/kimi-code/ - Kimi Code membership and tool path
- /value/kimi-access/ - Kimi access and promo-check guidance
- /value/smart-spend/ - Paid-stack routing strategy
- /compare/models/mid-range/ - Production spend-band comparison
Sources
- Z.AI GLM-5.2 docs
- Z.AI pricing
- Z.AI GLM Coding Plan overview
- Kimi K3 launch blog
- Kimi K3 quickstart
- Kimi K2.7 Code quickstart
- Kimi K2.7 Code pricing
- Kimi K2.7 Code release notes
- Kimi K2.6 quickstart
- Kimi K2.6 pricing
Last verified: July 18, 2026. Recheck API pricing, context limits, thinking-mode requirements, benchmark evidence, open-weight status, and subscription quota terms before committing production workloads.