October 2 routing correction: This page preserves its July evidence and recommendations. Z.AI now documents GLM-5.3; GLM-5.2 scores and prices below do not evaluate that revision. Historical status here does not imply API retirement. Use the selected model hub before choosing a new route.

Short verdict: Start with GLM-5.2 when the job needs 1M context, whole-repo context retention, or Z.AI-supported coding-tool routing at a lower token price. Start with Kimi K2.7 Code when you want cheaper Kimi coding economics, multimodal tool calls, or a HighSpeed API lane. Use Kimi K3 when the question is newest-Kimi or frontier-adjacent Kimi performance, and keep Kimi K2.6 in the comparison because it still has exact-match demand and remains relevant where non-thinking mode or existing Kimi integrations matter.

This is a shortlist guide, not a leaderboard. Vendor benchmark claims can help pick what to test, but the winning model is the one that fixes your actual bug with fewer retries and less human cleanup.

Quick Facts

ModelCurrent roleContextInput priceOutput price
GLM-5.2Z.AI July flagship comparison; now historical1M$1.40 / 1M, $0.26 cached$4.40 / 1M
Kimi K3Kimi newest flagship1M$0.30 cache hit / $3.00 cache miss$15.00 / 1M
Kimi K2.7 CodeCheaper Kimi coding release256K-class$0.19 cache hit / $0.95 cache miss$4.00 / 1M
Kimi K2.7 Code HighSpeedFaster K2.7 API lane256K-class$0.38 cache hit / $1.90 cache miss$8.00 / 1M
Kimi K2.6Prior Kimi multimodal/API comparison lane256K-class$0.16 cache hit / $0.95 cache miss$4.00 / 1M

Which Should You Test First?

If your constraint is…Start withWhy
Whole-repo context or very long tasksGLM-5.2Z.AI documents 1M context and 128K max output
Lowest cache-hit input costKimi K2.6 or Kimi K2.7 CodeKimi cache-hit input is lower than GLM’s public non-cached input anchor
Cheaper dedicated Kimi coding APIKimi K2.7 CodeK2.7 remains the lower-cost 256K coding lane; K3 is the newest flagship
Fast coding responsesKimi K2.7 Code HighSpeedSame K2.7 model, higher listed output speed, higher token prices
Supported-tool subscription economicsGLM Coding PlanZ.AI sells a GLM Coding Plan for supported coding tools
Non-thinking Kimi modeKimi K2.6K2.7 Code requires thinking; K2.6 can disable it
Multimodal coding inputsKimi K2.7 Code or K2.6Kimi docs emphasize text, image, and video input

Price And Context

GLM-5.2’s price case is not “cheapest raw token.” Its case is 1M context plus supported coding-tool economics. Kimi’s price case is cheaper input in cache-hit workflows and a simple OpenAI-compatible API path, with K2.7 Code HighSpeed available when speed is worth the higher output price.

Buyer questionGLM-5.2 answerKimi answer
“Which has bigger context?”GLM-5.2 at 1MKimi K2.7/K2.6 at 256K-class
“Which has cheaper output?”$4.40 / 1M$4.00 / 1M base K2.7/K2.6; $8.00 HighSpeed
“Which has cheaper cache-hit input?”$0.26 / 1M cached input$0.19 K2.7 Code, $0.16 K2.6
“Which has a subscription lane?”GLM Coding PlanKimi membership/Kimi Code paths, but verify current quota wording
“Which should handle routine work?”Test GLM if tool support fitsTest Kimi if API/tooling support fits

Tool Fit

WorkflowBetter first testReason
Claude Code-style supported-tool routingGLM-5.2 via Z.AIZ.AI documents Anthropic-compatible and supported-tool routes
OpenAI-compatible API experimentsBothBoth have OpenAI-compatible API paths
Kimi-native latest flagshipKimi K3 / Kimi CodeKimi-specific workflows should test K3 separately from cheaper K2.7 routing
Kimi-native budget codingKimi K2.7 Code / Kimi CodeKimi-specific workflows can keep K2.7 when price matters more than newest-model testing
OpenClaw/BYOK agent routingBothUse provider-specific docs and treat subscriptions separately from direct API billing
Very large repo auditGLM-5.21M context is the main spec advantage
Vision/video coding taskKimi K2.7 Code or K2.6Kimi docs explicitly emphasize multimodal input

Caveats

  • Z.AI benchmark language is vendor-published. Use it to decide what to test, not to claim “best model.”
  • Kimi K2.7 Code benchmark and speed language is also vendor-published. Measure your own cost per accepted patch.
  • Kimi K2.7 Code requires thinking mode. If your workflow needs non-thinking behavior, test K2.6.
  • Subscription quota and membership wording can move faster than API pricing docs. Treat checkout and current docs as final before purchase.
  • Context size is not quality. A smaller-context model can still win if it follows your repo conventions better.

Benchmark Evidence

benchmark artifact

GLM-5.2 vs Kimi Evidence

ModelProviderStatusContextInput priceOutput priceCoding signalTool-use signalBenchmark evidenceSpeedVerdictSourcesChecked
GLM-5.2Z.AI historical
Historical July comparison; Z.AI now documents GLM-5.3. These scores and prices do not evaluate 5.3. Historical status does not imply API retirement.
1M$1.40 / 1M
$1.40 input / $0.26 cached input / $4.40 output per 1M tokens
$4.40 / 1MZ.AI reports 62.1 on SWE-Bench Pro and 81.0 on Terminal-Bench 2.1.Supported-tool coding lane; BFCL score not imported.
  • Artificial Analysis Intelligence Index v4.1: 51 (independent)
  • SWE-Bench Pro: 62.1 (vendor)
  • Terminal-Bench 2.1: 81.0 (vendor)
Artificial Analysis flags higher output-token use; measure total cost per successful task.Dated July value candidate; retain for pinned integrations and historical comparisons. Check the newer GLM revision separately before new routing decisions.Z.AI GLM-5.3 successor documentation, Z.AI GLM-5.2 overview [archive], Z.AI pricing [archive], Artificial Analysis: GLM-5.2 article [archive], Artificial Analysis Intelligence Index v4.1, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Kimi K3Moonshot AI active
Kimi API/product flagship; Kimi K3 License weights. NVIDIA also lists K3 under trial terms as checked September 14; account access untested.
1M$3.00 / 1M
$0.30 cache-hit / $3.00 cache-miss input / $15.00 output per 1M tokens
$15.00 / 1MMoonshot reports strong max-reasoning launch-suite coding and agent results; AIHackers repo eval is not verified.Kimi API and Kimi Code support K3; Kimi docs warn to preserve full assistant history and avoid mid-session model switching.
  • Artificial Analysis Intelligence Index v4.1: 57 (independent)
  • Artificial Analysis output speed: 62 tokens/s (independent)
  • Moonshot launch benchmark suite: vendor-reported max-reasoning table (vendor)
  • AIHackers repo eval: not verified (site-owned)
Artificial Analysis reports 62 output tokens/s and flags high verbosity; measure total output cost per accepted task.Test as Kimi's newest 1M-context frontier-adjacent lane; keep K2.7 Code for cheaper routine Kimi coding until K3 passes local CAR tests.NVIDIA Kimi K3 model card, Kimi K3 launch blog [archive], Kimi K3 quickstart [archive], Kimi K3 API pricing [archive], Kimi K3 weights and license [archive], Kimi current model list [archive], Kimi Code model configuration, Artificial Analysis: Kimi K3 [archive]2026-08-01
Kimi K2.7 CodeMoonshot AI active
Cheaper routine Kimi coding API lane; HighSpeed is the same model at higher token prices.
256K$0.95 / 1M
$0.19 cache-hit / $0.95 cache-miss input / $4.00 output per 1M tokens; HighSpeed doubles those rates
$4.00 / 1MKimi K2.7 Code remains the lower-cost 256K coding lane after K3; independent normalized benchmarks are not imported.OpenAI-compatible API; thinking mode required in the documented K2.7 Code quickstart.
  • Program-Bench improvement vs K2.6: +10.4% (vendor)
  • MCP Mark Verified improvement vs K2.6: +11.4% (vendor)
  • SWE Marathon improvement vs K2.6: +76.2% (vendor)
  • Reasoning-token use vs K2.6: 30% lower (vendor)
  • AIHackers repo eval: not verified (site-owned)
HighSpeed model ID exists at a higher token price; latency not independently measured here.Cheaper routine Kimi coding API lane when Kimi routing fits and 256K context is enough.Kimi K2.7 Code quickstart [archive], Kimi K2.7 Code pricing [archive], Kimi Code K2.7 release notes [archive], SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Kimi K2.6Moonshot AI historical
Prior Kimi comparison lane retained for compatible integrations.
256K$0.95 / 1M
$0.16 cache-hit input / $0.95 cache-miss input / $4.00 output per 1M tokens
$4.00 / 1Mnot verifiedToolCalls and OpenAI-compatible API supported; BFCL score not imported.not verifiednot verifiedPrior Kimi comparison baseline where K2.6-specific integrations or multimodal behavior matter.Kimi models, Kimi K2.6 quickstart, Kimi K2.6 pricing, Artificial Analysis: Kimi K2.6, SWE-bench, Berkeley Function Calling Leaderboard2026-06-28
Claude Opus 4.8Anthropic historical
Historical comparison; use Opus 5.5 for current premium evaluation.
1M$5.00 / 1M
$5.00 input / $25.00 output per 1M tokens
$25.00 / 1MHistorical premium Claude baseline; use Opus 5 for new task-level comparisons.Still available for pinned integrations; new Claude premium routing should test Opus 5.
  • Artificial Analysis Intelligence Index v4.1: 56 (independent)
  • Artificial Analysis output speed: 57.3 tokens/s (independent)
Artificial Analysis measured 57.3 output tokens/s; provider and workload latency vary.Historical premium baseline. Use Claude Opus 5.5 for current Claude premium routing.Claude models overview [archive], Claude API pricing [archive], Artificial Analysis: Claude Opus 4.8 [archive], Artificial Analysis Intelligence Index v4.1, LMArena leaderboard dataset, Berkeley Function Calling Leaderboard2026-07-25

Kimi K3 has early independent evidence and vendor launch evidence. GLM-5.2 has an independent Artificial Analysis aggregate plus vendor coding scores. K2.7 results are vendor-reported relative deltas versus K2.6. Do not normalize them into one leaderboard.

Eval Checklist

Run the same tasks through each candidate:

EvalWhat to measureWinner signal
Bug fixOne real failing testSmallest correct patch, fewest retries
Refactor2-4 files with local style constraintsPreserves behavior and tests
Repo auditArchitecture map from docs and codeAccurate module boundaries
Tool callsMulti-step task with tool resultsKeeps reasoning/tool context intact
CostActual cache-hit/cache-miss and output usageLower cost per accepted result
ReviewPatch review against a known risky changeConcrete findings without invented policy

Sources


Last verified: July 18, 2026. Recheck API pricing, context limits, thinking-mode requirements, benchmark evidence, open-weight status, and subscription quota terms before committing production workloads.