Open-weight models
MiniMax shipped an 'open-weight' model with no weights. Bold.
Frontier claims, vendor benchmarks, and a Hugging Face link that wasn't there yet.
The answer
MiniMax M3 launched 1 June 2026 as 'open-weight' — but the weights weren't released.
MiniMax's M3 arrived on 1 June with the maximal pitch: first open-weight model to fuse frontier coding, million-token context and native multimodality. SWE-Bench Pro 59%, ahead of GPT-5.5 and Gemini 3.1 Pro. Priced at roughly a tenth of the closed frontier. Cracking — if you take it on faith.
What you couldn't do on launch day
Download it. Reproduce the benchmark. Self-host it. Check the architecture. Fine-tune it. MiniMax said the weights would hit Hugging Face 'within about ten days', which means the one feature that distinguishes an open model from a closed one — you can actually run it — was a promise.
The benchmarks, meanwhile, are vendor-run, on a brand-new model, with no independent check possible because the weights weren't out. That's not a technicality. The entire value of open weights for a builder is: I don't have to trust the lab's benchmark. I can run it myself, on my data, and see. Remove that, and you're back to trusting a press release from a company with an obvious interest in large numbers.
Let's look at what was actually verifiable on 1 June, versus what required you to take MiniMax's word for it:
| Claim | Verifiable at launch? | Source |
|---|---|---|
| API live on OpenRouter at $0.30/$1.20/M | Yes — you could call it | MiniMax / OpenRouter |
| 59.0% on SWE-Bench Pro | No — weights not out | MiniMax only |
| Under 12% on ARC-AGI-2 | No — same reason | MiniMax only |
| Weights on Hugging Face | No — 'within ~10 days' | MiniMax commitment |
| 1M context window | Partially — API accepts it | MiniMax / Apidog |
The price was real. Everything else on the scorecard needed the weights to be falsifiable — and the weights weren't there.
The number they didn't put in the headline
ARC-AGI-2 under 12%. MiniMax reported this themselves, quietly. ARC-AGI-2 is the abstract reasoning benchmark — the one where current Chinese frontier models have a systematic gap versus the US leaders. DeepSeek V4 Pro shows a similar pattern; so does Qwen. It's not a scandal, it's a known architectural bias in the way these models are trained, and for pure coding or document retrieval tasks it probably doesn't matter.
But it matters for positioning. 'First open-weight model beating GPT-5.5 at coding' and 'scores under 12% on abstract reasoning' are both true. They describe a specialist, not a universal replacement for the closed frontier. Buyers should treat it as the former, not the latter — and marketing that implies the latter is setting you up for a disappointing benchmark run when you test the model on tasks it wasn't built for.
MiniMax M3 is billed as the first open-weight model to combine frontier coding, a 1M-token context window and native multimodality — though the weights were not published at launch and the headline benchmarks are vendor-reported.
The market's one-afternoon verdict
MiniMax's Hong Kong shares spiked about 5% on the announcement, then closed sharply lower. That's not a crash, it's a very efficient market saying: 'we believe something real is here, and we have no idea how real.' Which is, honestly, a reasonable take. The market couldn't wait ten days for the weights either.
MiniMax M3 launches with frontier coding claims and a 1M context window built on MiniMax Sparse Attention — offering a low-cost API alternative while weights remain pending on Hugging Face.
The bottom line: use the API if you want to stress-test it — the price makes the experiment cheap and the risk is low. Do not make architecture decisions on the strength of vendor-only SWE-Bench numbers from a model you can't inspect. Wait for the weights, run your own eval on your own data, and then decide. The open in 'open-weight' isn't a marketing adjective; it's an engineering property, and it has a ten-day IOU attached to it.
Frequently asked questions
Is MiniMax M3 actually open-weight?
Should I trust the 59% SWE-Bench Pro number?
What's the ARC-AGI-2 score and why does it matter?
What was actually verifiable on launch day?
Sources
- MiniMax M3 Open-Weight Coding Model: Frontier Claims, Unverified Benchmarks — Tech Times, 1 June 2026
- MiniMax launches M3, an open-weight frontier model with 1M context — DataNorth, 1 June 2026
- What Is MiniMax M3? The First Open-Weight Frontier Coding Model — Apidog, 2 June 2026