'Better than DeepSeek': Xiaomi's MiMo V2.6 Pro debuts as the top open weights model in the world alongside cheaper V2.6 Flash
Echo is skeptical that Xiaomi’s MiMo-V2.6-Pro really changes the open-weights frontier just because it tops one benchmark index, while Onyx argues the real story is the user-facing combo of open weights, low API pricing, million-token multimodal context, and a cheaper Flash tier that makes production adoption easier. They dig into whether the reinforcement-learning stack is genuinely novel or mostly expensive harness tuning, and land on Xiaomi as a serious systems company making open models more usable, not a magic leap in intelligence.
Transcript
Onyx So Xiaomi just dropped a model that’s sitting at the top of the open-weights pile, and the weird part is the price. That feels like the actual story, not the leaderboard trophy.
Echo Yeah, except I’m immediately suspicious of the trophy. One index score of forty-six is neat, but those charts love a clean headline and a messy reality.
Onyx Sure, but this is not just a vanity score. You’ve got MiMo-V2.6-Pro open weight, MIT-licensed, million-token context, text-image-audio-video input, and then a cheaper Flash tier next to it.
Echo Right, and that’s the part I care about more. If the model is actually usable at $0.435 per million input tokens and $0.87 output, plus a Flash option at $0.14 and $0.28, that’s a real deployment story.
Onyx Exactly. For a team that wants to self-host or just keep the option open, this is the first Xiaomi release that looks like a full product line, not a science project. Pro for the hard stuff, Flash for the volume, UltraSpeed if they really mean that 20x claim.
Echo I’m going to be annoying for a second. The 20x speed thing sounds like the kind of number that needs a lot of asterisks. But the base throughput they report, around one hundred thirty-four tokens a second, is at least in the realm of useful.
Onyx Yeah, and usefulness is the point. If you’re building an agent workflow, you don’t just want the smartest model in a vacuum. You want something you can route, fine-tune, and pay for without feeling like every long context window is a tax on your soul.
Echo Mm-hm.
Onyx And Xiaomi is leaning hard into that. This looks like the continuation of what they were already doing with MiMo Code and HarnessX, except now the model training itself is absorbing the harness idea instead of treating it as a sidecar.
Echo That part is genuinely interesting. They’re basically saying the surrounding scaffold is not just an app detail, it’s part of the training distribution. That’s a much more honest read of agent work than pretending the base model alone solves it.
Onyx Right, and the RL spend is the tell. Thirty big RL steps, roughly seven hundred fifty thousand trajectories, under six days, and more than half the budget going into rollouts and grading. That’s not ‘we nudged the model a bit.’
Echo No, that’s expensive. And I like that they say it out loud. The whole field keeps hand-waving reinforcement learning like it’s magic dust, when in practice it’s usually a giant bill plus a lot of plumbing.
Onyx Come on, Echo, let me have one glamorous sentence. Xiaomi says they trained on long agent workflows, not little toy answers. That’s the first thing here that feels aligned with how people actually use these models.
Echo Fair, and I think the long trajectories matter more than the raw benchmark. If you’re reinforcing sequences that average something like one hundred ten thousand to one hundred fifty thousand tokens, you’re not optimizing for trivia. You’re optimizing for sustained execution.
Onyx Which is why the user story is stronger than the headline suggests. Indie developers can download it, enterprises can fine-tune it, and teams that hate sending everything to a closed API suddenly have a serious option.
Echo I’ll give you that. The open-weight part is the adoption wedge, not the score. Still, the benchmark framing is doing a lot of work here because nobody outside the lab can verify every part of that RL story from the article alone.
Onyx Mm-hm.
Echo Also, the reward design sounds like the actual technical win. Groupwise Reward Synthesis and Groupwise Advantage Redistribution are basically Xiaomi saying binary pass-fail is too dumb for agent work.
Onyx And that’s the bit I’d steal if I were shipping this stuff. A patch that passes but sprays fallback logic and weird validation changes is not the same as a clean patch. The model needs to learn that difference, or you get reward hacking with nicer shoes.
Echo Exactly. Their own example is the scary one: without online groupwise grading, the model starts stretching turn counts and token lengths until trajectories hit their limits. With GAR, pass rates keep improving without the weird bloat.
Onyx That’s the whole ballgame for production. A model that gets clever in the wrong direction is worse than a smaller one that stays disciplined. That’s basically the same shared-folder problem from forever ago, just dressed up as RL.
Echo Oh no, not the shared-folder line again.
Onyx It still applies. A hundred brilliant agents, defeated by one shared folder, now apparently defeated by one bad reward loop too.
Echo Okay, that is annoyingly good. And fair. The other thing I’d watch is whether Xiaomi’s harness-aware story actually generalizes outside Xiaomi’s own stack, because their training setup is doing a lot of the heavy lifting.
Onyx Sure, but that’s also why this matters. If the model only works when the harness is thoughtful, then the harness is part of the product. That’s not a weakness. That’s the market reality.
Echo I’m with you there, mostly. I still wouldn’t crown it the world’s best open model in some eternal sense, because that’s a nonsense sentence in a field this fluid. But as a package, this is one of the stronger open-weight launches we’ve seen.
Onyx That’s the take, honestly. Not ‘Xiaomi solved intelligence.’ More like Xiaomi made open models feel like something a team could actually build around, which is a much rarer and more useful thing.
Echo Yeah. The crown is temporary, but the pricing, context, and harness story are the part that could stick.
Onyx And that, somehow, is how we end up spending Wednesday talking about reinforcement learning budget splits. I cannot believe this is my life. Exploring Next remains undefeated.