This is the companion run to the GLM-5.3-Flash evaluation: same 113 tasks, same agent framework, same fixed task order, same binary grading. MiMo-V2.6-Flash-MOPD came out of it with 71 of 113 tasks resolved (62.8%) and 94.52% of all target tests passing, using markedly less effort per task than GLM — and failing in a way that is genuinely different.

The setup

   
Model MiMo-V2.6-Flash-MOPD (Xiaomi), W4A8 (MXFP4 experts + FP8 dense)
Server vLLM, tensor parallel 2 (2x RTX PRO 6000 Blackwell Max-Q)
Context 448,512 tokens, KV cache in fp8, DFlash speculative decoding
Agent mini-swe-agent, model_class: litellm, one task at a time
Per-task time budget 180 minutes, the corpus’s own limit
Task order Fixed seed, identical to the GLM run

Everything was served locally and quantized; the official board number for this model (about 68%) comes from a different setup with different weights, and I am not claiming to reproduce it.

The result

Metric Value
Tasks resolved 71 / 113 (62.8%), 95% CI [53.6 – 71.2]
Target tests passed (pooled) 5,555 / 5,877 = 94.52%
Target tests passed (mean per task) 90.59%
Regression tests passed (pooled) 231,316 / 231,352 = 99.984% (36 failures)
Regression tests passed (mean per task) 99.86%
Time per task median 15.8 min, mean 18.1 min, max 46.3 min
Time spent 31.4 h of agent time, 34.0 h of wall time
Effort 10,190 steps (90 per task), far fewer than GLM’s 139

The maximum is worth a sentence of its own: the longest task took 46 minutes, well under the corpus’s three-hour ceiling. Whatever else this model does, it does not sprawl. It reached the same wall-clock average as GLM (18.1 vs 19.0 minutes per task) with about a third fewer steps, which says its steps are longer — more reasoning or more generation packed into each turn.

Where the tasks land

Tasks by share of target tests passed

The shape is similar to GLM’s — most tasks either pass everything or land in the 90–99% band — but the tail is heavier. MiMo has 9 tasks below 50% of their target tests, against 4 for GLM, and 7 in the 50–89% band. When this model misses, it misses further.

By language

Language Resolved
Go 26/35
Python 21/34
TypeScript 20/34
Rust 1/5
JavaScript 3/5

Go and JavaScript are where MiMo beats GLM by one task each. Python and TypeScript are where it loses ground, and Rust is a rout in the other direction: MiMo resolved one of the five Rust tasks, GLM resolved three.

Where it earns its keep

Fourteen tasks were solved by MiMo and by nothing else in this comparison: awilix-async-container-initialization, bandit-interprocedural-taint-checks, go-git-worktree-merge-conflicts, happy-dom-deterministic-intersectionobserver, ipython-session-bundle-replay, koota-pair-relation-tracking, mashumaro-flattened-dataclass-fields, onedump-dump-encryption-pipeline, prometheus-typed-label-sorting, scc-bounded-memory-spilling, superjson-error-stack-serialization, tengo-callable-instance-isolation, testem-per-launcher-reports and updo-policy-alerting.

mashumaro-flattened-dataclass-fields is the cleanest example: MiMo scored 66/66 on the target tests, GLM scored 63/66. Three assertions, one point, and the two models are on opposite sides of the line. That is the sharpest illustration I have of how much the binary metric compresses.

Its failure modes

Of the 42 failures:

  • 23 near-misses in the 90–99% band.
  • 3 tasks with every target test passing and a regression costing the point: bandit-structured-nosec-directives, kombu-single-active-consumer-priority and vulture-persistent-analysis-cache — the last one is also in GLM’s list, so two different models broke the same existing behavior on the same task.
  • 7 in the 50–89% band and 9 below 50%, the tail I mentioned above.

MiMo against GLM, on the same 113 tasks

  MiMo-V2.6-Flash-MOPD GLM-5.3-Flash
Resolved 71/113 (62.8%) 76/113 (67.3%)
Pooled target tests 94.52% 97.38%
Mean per task 90.59% 94.71%
Pooled regression tests 99.984% (36 broken) 99.993% (17 broken)
Steps per task 90 139
Longest task 46 min 79 min
Tasks under 50% of target tests 9 4

GLM resolves more and gets closer on what it does not resolve; MiMo works more cheaply and has a narrower tail of catastrophic misses. Both statements are true at once, which is exactly the messiness the single pass-rate number hides — and the subject of the next post.

What this does not prove

  • One run per task, so the interval on 71/113 is [53.6, 71.2].
  • Quantized weights and a locally served endpoint, not the configuration behind the published number.
  • Five Rust tasks. The 1/5 here and the 3/5 for GLM are small-sample observations, not a ranking of the models on Rust.

Reproducibility

The run is driven by the same eval/ tooling described in the GLM post: a model registry entry with the serving metadata, a scheduler that survives being stopped, and one JSON record per batch plus a full agent trajectory and verifier log per trial. All 113 verifier logs for this model were audited for environment failures (network errors during dependency fetch, or a verifier that failed in both the base and new mode) and came back clean.