MiMo V2.6 Flash is a strong value candidate for coding, multimodal tasks and repeated agent work. Put it beside DeepSeek V4.1 Flash, GLM-5.3, and GPT-6 Luna on your shortlist. Its low listed API rates and released weights justify a trial; AIHackers has not measured its cost per accepted result.

Choose Flash first for a cost-focused evaluation, Pro when harder tasks justify more spend, and the separate 9B checkpoint for training research. Use the cost-saving playbook to decide whether the result is actually cheaper.

Quick facts

ModelIdentity and purpose
MiMo-V2.6-FlashAPI mimo-v2.6-flash; downloadable MiMo-V2.6-Flash-RL; efficiency-focused model with a 309B-total / 15B-active language backbone
MiMo-V2.6-ProAPI mimo-v2.6-pro; downloadable MiMo-V2.6-Pro-RL; flagship model with 1.02T total / 42B active parameters
Pro UltraSpeedAPI mimo-v2.6-pro-ultraspeed; separately priced speed option, not the budget default
Distill-Qwen-9BSeparate supervised-fine-tuned research starting point based on Qwen3.5-9B; not a small copy of Pro with the same capabilities

The Flash and Pro cards describe text, image, video and audio understanding. Xiaomi’s API model table lists 1M context and 128K maximum output, plus tool calls, structured output and web search. Context capacity does not guarantee accurate use of every token.

Which Xiaomi path should I choose?

Your jobAccess routeMain boundary
App, script or metered evaluationPay-as-you-go APIOrdinary API key and balance; tool charges can be additional
Regular use in supported agent toolsToken PlanDedicated credentials, shared package credits and supported-tool rules
Work inside MiMo DesktopDesktop membership or configured Token PlanTwo separate subscriptions; a Token Plan key can power Desktop without buying both
Run or adapt a model yourselfOpen weightsHardware, runtime support and component licenses still matter

Developer access paths

Start with Xiaomi’s API documentation and your tool’s provider setup. Xiaomi documents OpenAI- and Anthropic-compatible access, but a compatible protocol does not guarantee identical tool behavior. Record the exact model ID, tool version and billing route in your evaluation.

Plans and pricing

Overseas USD per million tokens, checked September 27, 2026. These are published rates, not an account-specific checkout quote. Xiaomi’s pricing table separates cached input, uncached input and output:

API modelCached inputUncached inputOutput
V2.6 Flash$0.0028$0.14$0.28
V2.6 Pro$0.0036$0.435$0.87
Pro UltraSpeed$0.036$4.35$8.70

Batch pricing is half the real-time token rates for Flash and Pro; UltraSpeed is excluded. Use Batch only when asynchronous completion suits the task. Web search is billed separately: the overseas list is $5 per 1,000 calls. Cache-hit pricing applies only to actual hits.

Token Plan: credits are not tokens

The current Token Plan table lists these ordinary monthly individual packages:

PlanMonthly priceShared credits
Lite$64.1 billion
Standard$1611 billion
Pro$5038 billion
Max$10082 billion

Standard, Pro and Max also have team plans billed per seat. Annual subscriptions and first-purchase offers have separate conditions; check renewal terms before committing.

The usage rules charge Flash 2 / 100 / 200 credits per cached-input / uncached-input / output token; Pro uses 2.5 / 300 / 600. For example, 10,000 uncached Flash input tokens plus 2,000 output tokens consume 1.4 million credits before any off-peak coefficient. That is an arithmetic illustration, not a task-capacity estimate.

All integrated tools draw from the package. Xiaomi lists a 0.8× consumption coefficient at 00:00–08:00 Beijing time. When credits run out, service pauses rather than automatically charging ordinary API balance. Switching to PAYG creates a different bill.

Desktop is a separate purchase decision

Xiaomi’s Desktop comparison says Token Plan users can configure a dedicated key and service address in Desktop. That does not activate Desktop membership. Token Plan currently excludes UltraSpeed; some higher Desktop membership tiers include it. Verify the app’s current price, quota and renewal display—those account-specific totals were not tested here.

Benchmarks and capability claims

The model cards publish these Xiaomi-reported release results. They are shortlist evidence, not an AIHackers head-to-head test:

EvaluationFlashProWhat to examine
DeepSWE v1.167.971.9Repository-task evaluation; exact checkpoint and agent setup matter
ProgramBench26.026.5A different programming task set; do not average it with DeepSWE
Toolathlon-Verified73.676.9Tool-using agent performance depends on the surrounding system

Source: release evaluation table. The compact card does not establish a normalized comparison for your provider, effort setting or harness. No independent reproduction or AIHackers V2.6 benchmark is claimed here. Historical V2.5 or connectivity tests do not evaluate V2.6.

The public RL dashboard shows checkpoints from a training run, with a different DeepSWE series. Do not splice those numbers into this release table. The dashboard walkthrough explains rewards, pass rates, grading and out-of-sample evaluation; benchmark literacy explains how to use the evidence.

What is open, and what can you run?

Xiaomi calls the September 22 release fully open source. The collection lists Pro-RL, Flash-RL and Distill-Qwen-9B with MIT model licenses. Xiaomi also describes released RL environments and training resources. This is more than API access, but it does not by itself prove that every pretraining input or the exact full training run can be reproduced.

Hardware reality: Flash’s 15B active parameters do not mean it fits like a dense 15B model. The full experts, cache and runtime need memory. Pro and Flash are substantial server deployments; the 9B SFT checkpoint is the smaller research entry point, not a tested laptop recommendation. Read open weights vs open source for the component map and concrete research links.

Privacy, terms, and data use

A model license is not a hosted-data agreement. Check Xiaomi’s current service and privacy terms against your data requirements before sending private work. This review did not establish account-specific retention, enterprise controls or regional access. Self-hosting also requires control of tools, telemetry and storage; see the local-inference guide.

Release chronology and migration

This guide originally covered V2.5 in June. V2.6 launched September 22, 2026. The deprecation notice schedules mimo-v2.5 and mimo-v2.5-pro to stop working at 10:00 Beijing time on October 21, 2026, with no automatic replacement. Change the configured name and test the integration before that date. API retirement does not remove previously downloaded weights.

Where Xiaomi fits

Keep MiMo as one useful value option. Flash earns a place in a small evaluation alongside DeepSeek, GLM and Luna; Pro is a separate escalation candidate. The deciding system is model + harness + tools + checks, and the cost includes failed attempts and review.

Evaluation plan

Use the same bounded bug fix, structured extraction and source-backed answer for each accessible candidate. Hold context, tools, retry budget and acceptance rules fixed. Record tokens, cache hits, tool charges, latency, accepted results and cleanup minutes. AIHackers MiMo V2.6 CAR: not-run.

Sources

Primary links are placed beside the claims above. Model identities and benchmark tables come from Xiaomi’s cards; prices and plan rules come from its public documentation. Checked September 27, 2026. Account access, checkout, serving performance and training reproduction remain untested.