Open weights, September 2026: close to the frontier, but not level with it
In July, when Moonshot released Kimi K3, Artificial Analysis reported that the gap between the best open-weight model and the best proprietary one was four points on its Intelligence Index, the smallest since February.
Today the best open-weight model is Xiaomi’s MiMo-V2.6-Pro at 46, and the best proprietary model is Claude Opus 5.5 at 58 (leaderboard, index v4.3.2). That’s a twelve-point gap.
Nothing got worse on the open side. What happened is the part of this market people forget: the closed labs ship too. Three things are true at the same time, and the decision you make depends on which one you’re looking at.
- Open weights follow the frontier by months, not years.
- The very top is still closed, and it gets cheaper with every release.
- The open models at the top of the table are not the ones most teams can run.
The gap moves in steps
Here’s the sequence on Artificial Analysis’ index over this year. The index gets re-versioned as its evaluations are replaced, so compare the gaps rather than the raw scores across rows:
| When | Best open weights | Best proprietary | Gap |
|---|---|---|---|
| April 30 | Kimi K2.6, MiMo V2.5 Pro (54) | GPT-5.5 (60) | 6 |
| July, Kimi K3 launch | Kimi K3 | — | 4 |
| September 7 | GLM-5.3, Kimi K3 (44) | Claude Fable 5.1, GPT-6 Astra (53) | 9 |
| September 26 | MiMo-V2.6-Pro (46) | Claude Opus 5.5 (58) | 12 |
The pattern is a sawtooth, not a trend line. An open release closes the gap, then a closed release reopens it. On September 22 Anthropic shipped Opus 5.5, and OpenAI shipped GPT-6 Sol and Luna; the same day Xiaomi’s MiMo-V2.6-Pro became the best open model there is. It tied xAI’s Grok 4.7, a closed model released the day before.
The honest summary is that open weights run about one release cycle behind. For most business work, one cycle behind is plenty. For the hardest agentic and reasoning work, it’s still a real difference.
Who ships the open frontier now
Almost every model near the top of the open table comes from a Chinese lab:
| Model | Lab | Size (total / active) | Context | License |
|---|---|---|---|---|
| MiMo-V2.6-Pro | Xiaomi | 1.02T / 42B | 1M | MIT |
| GLM-5.3 | Z.ai | 753B | 1M | GLM-5.3 License |
| Kimi K3 | Moonshot | 2.8T / 104B | 1M | Kimi K3 License |
| DeepSeek V4.1-Flash | DeepSeek | 552B / 8–16B | 1M | MIT |
| Qwen3.8-2.4T-A95B | Alibaba | 2.4T / 95B | 262K | Qwen3.8-Max License |
The Western open options sit in a different weight class. Mistral’s Small 4 is 119B with 6.5B active under Apache 2.0. Google’s Gemma 4 tops out at 31B. Meta, after more than a year without an open release, published Muse Glimmer in August: 30B, dense, Apache 2.0. Its larger Muse Spark stays closed. All of them are good models. None of them competes with the table above.
For a company that answers to a US compliance officer or a government customer, this matters. On September 8 the NSA, CISA and FBI issued advisory AA26-251A, naming DeepSeek, Moonshot, Alibaba, MiniMax, StepFun and Z.ai as having run “industrial-scale” distillation campaigns against US models. Two days later Anthropic’s threat report named seven China-based labs and about 190 million exchanges. The advisory is aimed at the labs, not at people who download their weights. The weights run on your hardware and send nothing anywhere. But “we run a model the federal government says was built on extracted data” is a sentence some procurement teams won’t sign, and you should find out whether yours is one of them before you pick a model, not after.
Read the license, not the headline
“Open weights” now covers everything from MIT to licenses with named thresholds. The three that matter this month:
- MIT (MiMo-V2.6-Pro, DeepSeek V4.1-Flash). Do what you want, keep the notice.
- Kimi K3 License. Internal use is exempt. If you sell the model as a service and that business passes $20M in revenue over 12 months, you need a separate agreement with Moonshot. Products above 100M monthly users or $20M a month must display “Kimi K3” in the interface.
- GLM-5.3 License. GLM-5.2 was MIT; 5.3 is not. The restriction only applies to operators with over $10B in revenue who sell inference as a service, so it’s aimed at hyperscalers. For a law firm running it internally, nothing changes. But it’s a license that changed between point releases, and the next one can too.
Alibaba’s split is the one to watch. Qwen3.8-27B is Apache 2.0. The Max-class open checkpoint, Qwen3.8-2.4T-A95B, ships under its own license, text-only, without the vision, 1M-token context and built-in tools that the paid Qwen3.8-Max API has. That’s the open model trailing the lab’s own API by design, not by accident.
And none of these are open source in the sense the OSI means it. You get weights, not the training data or the recipe. You can run it and fine-tune it; you can’t audit what went into it.
Where the cloud is winning
The cloud’s advantage used to be capability. Now it’s capability and price.
| Model | Input / output per 1M tokens |
|---|---|
| GPT-6 Luna | $0.10 / $0.50 |
| DeepSeek V4.1-Flash (DeepSeek API, peak) | $0.30 / $1.20 |
| Gemini 3.8 Flash (to Dec 31) | $0.75 / $3.75 |
| GPT-6 Sol | $2 / $10 |
| Claude Opus 5.5 | $4 / $20 |
Prices as of September 26, 2026. They change monthly.
Look at the first two rows. A closed model from OpenAI now costs less per token at list price than DeepSeek’s own API for its open model (DeepSeek pricing). Claude Opus 5.5, the model at the top of the index, came in 20% cheaper than the Opus it replaced (Anthropic pricing), and GPT-6 Sol and Luna came in at half the price of GPT-5.6 (VentureBeat). Gemini 3.8 Flash is the exception worth noting: its price doubles on January 1.
This makes the argument I made in the local-vs-cloud cost post stronger, not weaker. If the only reason to run open weights is that tokens are cheaper, that reason keeps getting smaller.
The middle option: open weights in someone else’s cloud
“Open” no longer means “on your own hardware.” On September 18, Kimi K3 became generally available on Amazon Bedrock, with AWS stating that data stays inside its boundary and isn’t shared with Moonshot. You get a frontier-class open model with no GPUs to buy, under your existing cloud agreement.
That’s the right answer more often than people expect. It keeps you off any single lab’s API, and if you later have to move on-prem, you move the same weights, not a different model with different behavior.
What actually fits in your server room
The top of the open table is a datacenter conversation. Kimi K3’s weights alone are about 1.5 TB. MiMo-V2.6-Pro at 4-bit is over 500 GB. These are 8-to-16-GPU models. The mixture-of-experts design makes them fast for their size, but, as I covered in what one GPU can run, it doesn’t make them small: every expert has to be in memory.
What a firm realistically deploys is one or two tiers down:
- One 24–48 GB card: Qwen3.8-27B (dense, Apache 2.0, 262K context), Gemma 4 31B, Muse Glimmer 30B. Around 16–18 GB at 4-bit.
- 128 GB unified memory (DGX Spark, Strix Halo) or two datacenter cards: Mistral Small 4 at 4-bit, around 65–75 GB, with only 6.5B active, so it’s quick. See the DGX Spark vs Ryzen AI Max+ 395 comparison.
- A rented or owned 8-GPU node: DeepSeek V4.1-Flash and up.
The good news is how much this tier improved this year. For document Q&A, extraction and drafting over your own files, where the retrieval does most of the work, it’s enough.
So what do I recommend?
- If your data can leave the building: use the cloud. Pick by price and capability for the task. Opus 5.5 or GPT-6 Sol for the hard work, Luna or Flash for the volume.
- If you want open weights but not hardware: run them on Bedrock or a similar managed service. You keep the option to bring the same model in-house later.
- If the data can’t leave: a 27B–120B open model on your own hardware, with good retrieval in front of it, covers most real work. Accept that you’re a cycle behind on the hardest problems, and route those to the cloud where the data allows.
- Whichever you choose: read the license of the exact checkpoint, ask your compliance team about model origin before you standardize, and treat every score in this post as a snapshot. The closed labs ship, the open labs catch up, and the gap will have moved again by the time you read this.