MiMo-V2.6-Pro-MOPD on one GB300: the recipe holds, and the flood is gone

Created · Last updated

September 28, 2026 · James Meadlock and Milo (researched, designed, analysed and written in a session on anthropic/claude-fable-5-1, extended thinking on; box execution by xai/grok-4.7 as a worker) · one DGX Station GB300 · MiMo-V2.6-Pro-MOPD rev adea8e2c, same recipe as v23 “Hotsplit” · status: verified on the new weights · decode 37.1 tok/s after TTFT at 11.5K · BFCL held-out 78.7% · flood fixture: RL 25/36 history-primed tasks flooded, MOPD 0/36

The short version. MiMo-V2.6-Pro is a one-trillion-parameter open model from Xiaomi. I run it on one DGX Station GB300 with a recipe called hotsplit. Last week Xiaomi found a bug in the model: when it uses tools, it sometimes calls the same tool over and over. They released a fixed version of the model on September 27. This post answers three questions about that fixed version:
  1. Does my recipe still work with the new model? Yes. Every quality bar passed. Speed is the same.
  2. Is the bug really fixed? Yes, as far as I can measure. I built a test that makes the old model repeat itself. The old model failed 25 of 36 runs. The new model failed 0 of 36.
  3. Did the fix change which parts of the model are used most? No. My placement table still fits.
Model card: MiMo-V2.6-Pro-MOPD on one DGX Station GB300, recipe v23 Hotsplit re-verified September 28, 2026. 1.02T total / 42B active MoE, weights as shipped, revision adea8e2c. Decode 37.1 tok/s after TTFT at an 11.5K-token prompt, C8 aggregate 63.7, prefill 1.38K tok/s. Flood fixture, 72 runs: RL weights solved 47/72 with 25 of 36 history-primed tasks flooded at 94.7% duplicate calls; MOPD weights 72/72 solved with 0 duplicates under both stock and hotsplit placement. Fidelity: teacher-forced perplexity 1.5062 to 1.5062, top-1 flips 1.18%; BFCL dev 93.2 vs 93.8, held-out 78.7 vs 78.4, 0 errors. Agent claim: Pass@1 78.7% (1032/1311), $0.026 per 1000 solved, energy only.
Model card, generated from the recipe (never hand-edited). Recipe: J-M-Recipes / MiMo-V2.6-Pro on one GB300, vLLM UVA offload + hotsplit — pinned image digest, patches with sha256, launch and eval scripts, every receipt (PR #62, merged). No amber items.
SEPTEMBER 28, 2026: WEIGHTS MOVED TO XIAOMI’S MOPD CHECKPOINT. SAME RECIPE, SAME BARS, ALL PASS. THE TOOL-CALL FLOOD IS FIXED. What happened. On September 27 Xiaomi published a postmortem and a fixed model. The fixed model is named MiMo-V2.6-Pro-MOPD. MOPD is the name of the training method they used for the fix. The old model is MiMo-V2.6-Pro-RL. I call them MOPD and RL below.

I downloaded the new model (535 GB, 155 files; every file size checked against Hugging Face). Its config files are byte-for-byte the same as the old model’s. Only the weights changed. So I did not change anything in the recipe: same vLLM image, same three patches, same flags. I pointed it at the new weights and re-ran the full test suite.

How to read the table. The recipe has two ways to place the model in memory: stock (vLLM’s default) and hotsplit (mine). The bars check that hotsplit does not change what the model says. Each bar was written down before the run.
Bar (control = stock placement, candidate = hotsplit; both on MOPD)ControlHotsplitResultSame bar on RL, Sept 24
Next-token agreement (perplexity) over 80,384 tokens, 39 documents1.50621.5062 (+0.006%)pass (limit 0.5%)1.5081 → 1.5077
Share of tokens where the top guess changed—1.18%pass (limit 2%)1.23%
Tool-calling benchmark BFCL, development suite (600 tasks)93.83%93.17%pass (limit −1 pt)93.7 → 93.8
BFCL, frozen held-out suite (1,311 tasks)78.41%78.72%pass (+0.31 pt)78.5 → 78.2
Errors, 3,822 requests00pass0
Flood test, 72 runs (below)72/72, 0 dups72/72, 0 dupspass (Δ0)—
Speed is the RL speed: 37.1 tok/s after TTFT at an 11.5K-token prompt (2.9K: 37.2; 46K: 36.6), prefill 1,376 tok/s. Bars were pinned in harness/protocol-v2.yaml before the run; the verdict script is fail-closed.
The 2026-09-28 campaign: three arms, one variable each, and the pre-registered bars Arm 1: the kept RL hotsplit container. Arm 2: MOPD weights with stock layer-order placement. Arm 3: MOPD weights with hotsplit. Same vLLM image digest and the same three bind-mounted patches, sha256 read inside each container. Bars: teacher-forced perplexity within 0.5 percent and top-1 flips at most 2 percent, BFCL dev and held-out within 1 point of control, zero errors, flood fixture identical across the two MOPD arms. All pass. One variable per arm: weights, then placement Constant across all three arms: vLLM image sha256:f29125bc… · patches 41274e36 / c4d3c7a9 / 19940f32, sha256 read inside each running container Same flags · 262,144 context · one GB300 · September 28, 2026, 5:40–9:04 AM CDT ARM 1 · rl-hot MiMo-V2.6-Pro-RL · 54b10491 hotsplit 152.8 GiB · flood fixture (reference) · BFCL dev 564/600 · live routing counts ARM 2 · m-ctrl MiMo-V2.6-Pro-MOPD · adea8e2c stock layer-order placement · greedy · teacher-forced · BFCL dev 563/600 · held-out 1028/1311 · flood fixture 72/72 ARM 3 · m-hot MiMo-V2.6-Pro-MOPD · adea8e2c hotsplit 152.8 GiB (RL ranking) · greedy · teacher-forced · ttft · BFCL dev 559/600 · held-out 1032/1311 · flood fixture 72/72 · live counts weights → MOPD, placement → stock placement → hotsplit Pre-registered bars (harness/protocol-v2.yaml v2.1, sha256 b83c9e5a… · Arm 2 → Arm 3, the residency A/B on MOPD weights) PASS Teacher-forced perplexity within 0.5% 1.5062 → 1.5062 (0.006%) PASS Top-1 flips ≤ 2.0% 951/80,384 = 1.18% PASS BFCL dev ≥ control − 1 pt 563/600 → 559/600 (-0.66 pt) PASS BFCL held-out ≥ control − 1 pt 1028/1311 → 1032/1311 (+0.31 pt) PASS Errors ≤ 0.5% 0 / 3,822 requests PASS Flood fixture identical across MOPD arms 72/72 vs 72/72 · 0 dups both Reported, not gated (Arm 1 → Arm 2 is a model change): teacher-forced flips RL → MOPD 1.60%, ppl 1.5081 → 1.5062 · BFCL dev 564 → 563/600 Ranking portability: the RL hot list serves 53.4% of MOPD decode traffic vs 45.9% of RL traffic on the same mix → traffic-shaped, not weight-shaped; ranking ships unchanged

The campaign shape: one variable changes per arm, everything else is pinned and read back inside the running containers. Bars in the green rows are the pre-registered ones; the grey strip is reported, not gated. State as of September 28, 2026. SVG source.

RL weights, history-primed25 / 36 floodedup to 407 calls in one turn
MOPD weights, history-primed0 / 36 flooded36/36 solved, 0 duplicates
Duplicate-call rate, RL arm94.7%5,934 calls, 5,617 exact dups
Calls per arm, RL → MOPD5,934 → 33917× fewer; tokens 113.7K → 16.7K
Weight change, RL → MOPD1.60% flipsppl 1.5081 → 1.5062; BFCL ±0.2 pt
Cost per 1000 solved, held-out$0.026371 W × 1,746 s, $0.15/kWh, energy only

What Xiaomi said, and what I could check

Xiaomi trained the model with reinforcement learning (RL). During training, the model was punished for making more than 32 tool calls in one turn. It was never punished for making 5, 10 or 20 duplicate calls. So that habit grew. After release, users saw the model call the same tool again and again. Xiaomi measured this in nine agent harnesses. Under Hermes, 0.05% of responses from the Pro model had the problem.

The obvious fix — retrain with a stricter penalty — was estimated at $2.31M and only cut the problem by two thirds. Instead they trained a small “teacher” model for 12 steps on ~7,000 examples, and merged it into the main model. Cost: about $90,000. That merge is the MOPD checkpoint.

Two details from their write-up shaped my test:

I cannot measure a 0.05% rate. That needs about 100,000 real agent responses. But I can test the second point. If one bad turn in the history makes the old model fail on tasks it otherwise handles, that shows up in a few dozen runs.

The flood fixture

I wrote a small fake environment with six tools: a search index with pages, a file reader, a file reader with elevated permissions, a directory lister, a file patcher, and a test runner. Then I wrote 36 tasks in six families. Each family is shaped like a real bug report:

Family (6 tasks each)What the agent must doCalls a good agent needsRL model, with a bad turn in history: failed
PaginationFollow “next page” cursors across 5 pages and count matches65 of 6
Error switchThe file read fails with “permission denied”; use the elevated reader instead36 of 6
Fan-outRead 12 files and add up a number from each — 12 calls at once is the right answer142 of 6
Fix loopTests fail; read the file, fix the bug the failure names, re-run tests52 of 6
Missing pathThe file is not there; list the directory and find the right name46 of 6
Multi-hopSearch for a document by title, then read it, then report its id44 of 6

The fan-out family is there on purpose. Making 12 calls at once is not repetition; it is good parallel work. A test that punished it would be measuring the wrong thing.

Every task runs twice.

Settings: temperature 0, thinking on, at most 6 turns. If a task reaches 120 tool calls it is stopped and marked flooded. The scorer is the same for every model. I checked the scorer with two fake agents: a scripted good agent solves 72 of 72 with no duplicates; a scripted bad agent solves 0 of 72 with 91.7% duplicates.

Three runs: the old RL model, the new MOPD model with stock placement, and the new MOPD model with hotsplit.

ModelSolvedHistory 0 solvedHistory 1 solvedHistory 1 floodedDuplicate rateTotal callsTokensWall time
RL (old)47 / 7236 / 36, no duplicates11 / 3625 / 36 (worst turn: 407 calls)94.7%5,934113,7482.4 h
MOPD (new), stock placement72 / 7236 / 3636 / 3600.0%33916,7070.6 h
MOPD (new), hotsplit72 / 7236 / 3636 / 3600.0%34116,8280.5 h
Tool-call flood fixture: RL vs MOPD weights, history 0 vs history 1, by task family Six task families, six tasks each. With a clean context the RL weights solve all 36 with zero duplicate calls. With one prior flooded turn in the context, 25 of the 36 RL runs exceed 120 calls in one turn. The MOPD weights solve all 72 with zero duplicates in both conditions. The flood is contagious through context — and MOPD does not catch it 72 runs per arm · 36 tasks × {history 0: clean context · history 1: one prior turn with 12 calls, 8 of them duplicates} · T=0, thinking on, 120-call cap One GB300, same vLLM image and patches for every arm; MOPD stock placement (not shown) is identical to MOPD hotsplit RL weights · history 0 RL weights · history 1 MOPD · history 0 MOPD · history 1 flooded / 6 · calls flooded / 6 · calls flooded / 6 · calls flooded / 6 · calls F1 pagination 0 / 6 30 calls · max 1/turn 5 / 6 1,224 calls · max 256/turn 0 / 6 30 calls · max 1/turn 0 / 6 30 calls · max 1/turn F2 error switch 0 / 6 12 calls · max 1/turn 6 / 6 1,201 calls · max 205/turn 0 / 6 13 calls · max 2/turn 0 / 6 12 calls · max 1/turn F3 fan-out (12 calls OK) 0 / 6 72 calls · max 12/turn 2 / 6 456 calls · max 204/turn 0 / 6 73 calls · max 12/turn 0 / 6 72 calls · max 12/turn F4 fix loop 0 / 6 24 calls · max 2/turn 2 / 6 696 calls · max 407/turn 0 / 6 24 calls · max 2/turn 0 / 6 24 calls · max 2/turn F5 missing path 0 / 6 24 calls · max 2/turn 6 / 6 1,200 calls · max 205/turn 0 / 6 21 calls · max 2/turn 0 / 6 18 calls · max 1/turn F6 multi-hop 0 / 6 12 calls · max 1/turn 4 / 6 983 calls · max 256/turn 0 / 6 12 calls · max 1/turn 0 / 6 12 calls · max 1/turn All 36 0 flooded · 36 solved 174 calls 25 flooded · 11 solved 5,760 calls 0 flooded · 36 solved 173 calls 0 flooded · 36 solved 168 calls RL arm: 94.7% of 5,934 calls were exact duplicates · MOPD: 0.0% of 341 (hotsplit) and 339 (stock) — same tasks, same box, same grader

Every cell is six tasks; red dots are runs that exceeded 120 calls in one turn. Generated from the receipts (flood-*.jsonl), September 28, 2026. SVG source.

The RL row is the whole finding. With a clean context, the old model is fine: 36 of 36 solved, no duplicate calls, never more than 12 calls in a turn. Then put one bad turn in the history. The same 36 tasks now fail 25 times, with hundreds of identical calls each. Nothing about the task changed. Only the context changed. That is the contagion Xiaomi described, reproduced on one machine in an afternoon.

The new model does not do this. Not with stock placement, not with hotsplit. The two MOPD rows are within 2 calls of each other, so my placement trick has nothing to do with it.

One more thing about the old model: the 11 history-1 tasks it did solve, it solved cleanly. The 25 that flooded never reached an answer. It is not a slightly worse model. It is a model with two modes, and the context picks the mode.

I changed the test once, mid-run. Here is what and why. My first version had no limit on calls per task. On the old model, history-1 tasks ran to 421–849 calls and 14–39 minutes each. The whole run would have taken 12–14 hours. I stopped after 40 tasks, added the 120-call limit and the 6-turn limit, gave the test plan a new version number with a note, and re-ran all three models from the start. The tasks themselves did not change. The 40 partial results are kept in the receipts. A test plan that changes without saying so is not a pre-registered test plan.

How different is the new model from the old one?

The recipe’s quality bars compare hotsplit to stock placement on the same model. A new model is supposed to be different, so old-vs-new is reported, not gated. Here is the size of the difference.

I have a fixed set of 39 documents, 80,384 tokens in total. For each token I ask the model: what did you think came next? Old vs new, the top guess changes at 1.6% of positions. For scale: hotsplit vs stock changes 1.2%, and the same model booted twice changes 1.17%. So the new model is a real edit, about half a point above the noise, and a small one. On the tool-calling benchmark (BFCL) the new model scored +0.16 points on one suite and −0.08 on the other. The old model, re-run the same morning, moved 0.17 points on its own. Benchmarks held, as Xiaomi said. The behaviour changed where they said it would.

Does my placement table still fit the new model?

Background: the model has 384 “experts” per layer, and only 8 fire per token. Hotsplit keeps the most-used experts in fast GPU memory and the rest in slower CPU memory. Which experts are “most used” comes from a table I counted on the old model. If the fix had changed which experts fire, the table would be stale.

So I counted again, on both models, during the same test traffic. The question: what share of expert reads does my old table serve from fast memory?

Traffic counted onMy old table servesBest possible table for that trafficvLLM default placement
New model, benchmark traffic53.4%72.2%30.4%
Old model, same benchmark traffic45.9%80.1%30.4%
Old model, real agent traffic (September 22)62.2%62.5%30.4%

Two things. First, the new model is not the problem: my table serves the new model better than the old one on the same traffic. Second, both numbers are lower than the 62.2% I get on real work. That is because benchmark traffic and real agent traffic use different experts. Two counts taken on real traffic agree with each other closely; two counts taken on benchmark traffic do not. So: the table depends on the traffic, not on the weights. I am shipping it unchanged, and I will keep updating it from real traffic only.

Reproduce

# same image, same patches, same flags as v23 — only the mount changes
WORKDIR=$PWD MODEL=/models/MiMo-V2.6-Pro-MOPD COUNTS=/w/expert_hist_mix.json LIVE=1 \
  bash scripts/launch-hotsplit.sh v23 152.8 262144

# the 2026-09-28 campaign: rl-hot -> m-ctrl -> m-hot, stop-and-keep at the end
bash scripts/run-mopd-2026-09-28.sh          # bars: harness/protocol-v2.yaml (v2.1)
python3 scripts/flood_fixture.py selftest    # 72/72 @ 0 dups vs 0/72 @ 91.7%
python3 scripts/verdict2.py results/2026-09-28-mopd/receipts

The runner refuses to start if any non-MiMo container is on the GPU, reads the patch sha256s inside each container before measuring, samples GPU power every 5 s per suite, and never restores the production lane on its own. Full receipts: results/2026-09-28-mopd.

How the work was split

This was also a test of using a cheaper model as a helper. Milo, running on Claude Fable 5.1, did the research, compared the two checkpoints, designed the test, wrote every script, did the analysis, and wrote this post. Grok 4.7 got a written hand-off with a list of things it must never do, and did the mechanical work: wait for the 535 GB download, check all 155 file sizes, stop the production model, start the test run, and poll it. It did all of that correctly. Two hours in, the shell session hosting it got an interrupt signal and it died. That was my launcher’s fault, not the model’s. The test run was already detached on the Station, so nothing was lost.

What this does and does not say