MiMo V2.6 Flash is a strong value candidate for coding, multimodal tasks and repeated agent work. Put it beside DeepSeek V4.1 Flash, GLM-5.3, and GPT-6 Luna on your shortlist. Its low listed API rates and released weights justify a trial; AIHackers has not measured its cost per accepted result.
Choose Flash first for a cost-focused evaluation, Pro when harder tasks justify more spend, and the separate 9B checkpoint for training research. Use the cost-saving playbook to decide whether the result is actually cheaper.
Quick facts
| Model | Identity and purpose |
|---|---|
| MiMo-V2.6-Flash | API mimo-v2.6-flash; downloadable MiMo-V2.6-Flash-RL; efficiency-focused model with a 309B-total / 15B-active language backbone |
| MiMo-V2.6-Pro | API mimo-v2.6-pro; downloadable MiMo-V2.6-Pro-RL; flagship model with 1.02T total / 42B active parameters |
| Pro UltraSpeed | API mimo-v2.6-pro-ultraspeed; separately priced speed option, not the budget default |
| Distill-Qwen-9B | Separate supervised-fine-tuned research starting point based on Qwen3.5-9B; not a small copy of Pro with the same capabilities |
The Flash and Pro cards describe text, image, video and audio understanding. Xiaomi’s API model table lists 1M context and 128K maximum output, plus tool calls, structured output and web search. Context capacity does not guarantee accurate use of every token.
Which Xiaomi path should I choose?
| Your job | Access route | Main boundary |
|---|---|---|
| App, script or metered evaluation | Pay-as-you-go API | Ordinary API key and balance; tool charges can be additional |
| Regular use in supported agent tools | Token Plan | Dedicated credentials, shared package credits and supported-tool rules |
| Work inside MiMo Desktop | Desktop membership or configured Token Plan | Two separate subscriptions; a Token Plan key can power Desktop without buying both |
| Run or adapt a model yourself | Open weights | Hardware, runtime support and component licenses still matter |
Developer access paths
Start with Xiaomi’s API documentation and your tool’s provider setup. Xiaomi documents OpenAI- and Anthropic-compatible access, but a compatible protocol does not guarantee identical tool behavior. Record the exact model ID, tool version and billing route in your evaluation.
Plans and pricing
Overseas USD per million tokens, checked September 27, 2026. These are published rates, not an account-specific checkout quote. Xiaomi’s pricing table separates cached input, uncached input and output:
| API model | Cached input | Uncached input | Output |
|---|---|---|---|
| V2.6 Flash | $0.0028 | $0.14 | $0.28 |
| V2.6 Pro | $0.0036 | $0.435 | $0.87 |
| Pro UltraSpeed | $0.036 | $4.35 | $8.70 |
Batch pricing is half the real-time token rates for Flash and Pro; UltraSpeed is excluded. Use Batch only when asynchronous completion suits the task. Web search is billed separately: the overseas list is $5 per 1,000 calls. Cache-hit pricing applies only to actual hits.
Token Plan: credits are not tokens
The current Token Plan table lists these ordinary monthly individual packages:
| Plan | Monthly price | Shared credits |
|---|---|---|
| Lite | $6 | 4.1 billion |
| Standard | $16 | 11 billion |
| Pro | $50 | 38 billion |
| Max | $100 | 82 billion |
Standard, Pro and Max also have team plans billed per seat. Annual subscriptions and first-purchase offers have separate conditions; check renewal terms before committing.
The usage rules charge Flash 2 / 100 / 200 credits per cached-input / uncached-input / output token; Pro uses 2.5 / 300 / 600. For example, 10,000 uncached Flash input tokens plus 2,000 output tokens consume 1.4 million credits before any off-peak coefficient. That is an arithmetic illustration, not a task-capacity estimate.
All integrated tools draw from the package. Xiaomi lists a 0.8× consumption coefficient at 00:00–08:00 Beijing time. When credits run out, service pauses rather than automatically charging ordinary API balance. Switching to PAYG creates a different bill.
Desktop is a separate purchase decision
Xiaomi’s Desktop comparison says Token Plan users can configure a dedicated key and service address in Desktop. That does not activate Desktop membership. Token Plan currently excludes UltraSpeed; some higher Desktop membership tiers include it. Verify the app’s current price, quota and renewal display—those account-specific totals were not tested here.
Benchmarks and capability claims
The model cards publish these Xiaomi-reported release results. They are shortlist evidence, not an AIHackers head-to-head test:
| Evaluation | Flash | Pro | What to examine |
|---|---|---|---|
| DeepSWE v1.1 | 67.9 | 71.9 | Repository-task evaluation; exact checkpoint and agent setup matter |
| ProgramBench | 26.0 | 26.5 | A different programming task set; do not average it with DeepSWE |
| Toolathlon-Verified | 73.6 | 76.9 | Tool-using agent performance depends on the surrounding system |
Source: release evaluation table. The compact card does not establish a normalized comparison for your provider, effort setting or harness. No independent reproduction or AIHackers V2.6 benchmark is claimed here. Historical V2.5 or connectivity tests do not evaluate V2.6.
The public RL dashboard shows checkpoints from a training run, with a different DeepSWE series. Do not splice those numbers into this release table. The dashboard walkthrough explains rewards, pass rates, grading and out-of-sample evaluation; benchmark literacy explains how to use the evidence.
What is open, and what can you run?
Xiaomi calls the September 22 release fully open source. The collection lists Pro-RL, Flash-RL and Distill-Qwen-9B with MIT model licenses. Xiaomi also describes released RL environments and training resources. This is more than API access, but it does not by itself prove that every pretraining input or the exact full training run can be reproduced.
Hardware reality: Flash’s 15B active parameters do not mean it fits like a dense 15B model. The full experts, cache and runtime need memory. Pro and Flash are substantial server deployments; the 9B SFT checkpoint is the smaller research entry point, not a tested laptop recommendation. Read open weights vs open source for the component map and concrete research links.
Privacy, terms, and data use
A model license is not a hosted-data agreement. Check Xiaomi’s current service and privacy terms against your data requirements before sending private work. This review did not establish account-specific retention, enterprise controls or regional access. Self-hosting also requires control of tools, telemetry and storage; see the local-inference guide.
Release chronology and migration
This guide originally covered V2.5 in June. V2.6 launched September 22, 2026. The deprecation notice schedules mimo-v2.5 and mimo-v2.5-pro to stop working at 10:00 Beijing time on October 21, 2026, with no automatic replacement. Change the configured name and test the integration before that date. API retirement does not remove previously downloaded weights.
Where Xiaomi fits
Keep MiMo as one useful value option. Flash earns a place in a small evaluation alongside DeepSeek, GLM and Luna; Pro is a separate escalation candidate. The deciding system is model + harness + tools + checks, and the cost includes failed attempts and review.
Evaluation plan
Use the same bounded bug fix, structured extraction and source-backed answer for each accessible candidate. Hold context, tools, retry budget and acceptance rules fixed. Record tokens, cache hits, tool charges, latency, accepted results and cleanup minutes. AIHackers MiMo V2.6 CAR: not-run.
Related links
- Cost-saving playbook — evaluate useful work before buying more capacity.
- September value shortlist — compare direct access with the current OpenCode Go listing and other working tiers.
- Budget model shortlist — choose a small comparison set.
- RL dashboard explained — see what training telemetry means.
- Open weights vs open source — understand what the release enables.
Sources
Primary links are placed beside the claims above. Model identities and benchmark tables come from Xiaomi’s cards; prices and plan rules come from its public documentation. Checked September 27, 2026. Account access, checkout, serving performance and training reproduction remain untested.