Kimi K3 and Qwen 3.8 prove open models can reach the frontier. Why the economics of frontier labs favor infrastructure owners — and why Anthropic's model-only position risks unravelling.
There was never a question of if open models can match a frontier. Open just means you released the weights. The question is how they trained it. Most of the models coming out of China are distilled from larger commercial frontier models, and those are not the same thing. More situationally brittle.
Can be really good in narrow lanes, but not the same category at all. And that doesn’t show up in benchmarks.
People really need to stop parroting this line uncritically. The process takes time because even when you’re distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn’t been much time to do that. However, models like Kimi also do better than Fable or GPT on a lot of tasks, which means it’s not just distillation but also difference in architecture. You can watch this talk from Kimi founder to see how Kimi was actually trained and why it performs well.
It’s also absolutely hilarious that people think only Chinese companies use distillation, as if Anthropic or OpenAI are above that or something. Not to mention that they basically ignored copyrights on all the data the siphoned and are now crying that people aren’t respecting their terms of use.
The reality is that China tops the world in artificial intelligence publications today. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. Anybody who thinks China is simply distilling glorius American models is not engaging with reality.
Not to mention that US companies models constantly distill each other as Musk was forced to admit under oath. This whole narrative has just been a massive cope.
There was never a question of if open models can match a frontier. Open just means you released the weights. The question is how they trained it. Most of the models coming out of China are distilled from larger commercial frontier models, and those are not the same thing. More situationally brittle.
Can be really good in narrow lanes, but not the same category at all. And that doesn’t show up in benchmarks.
People really need to stop parroting this line uncritically. The process takes time because even when you’re distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn’t been much time to do that. However, models like Kimi also do better than Fable or GPT on a lot of tasks, which means it’s not just distillation but also difference in architecture. You can watch this talk from Kimi founder to see how Kimi was actually trained and why it performs well.
It’s also absolutely hilarious that people think only Chinese companies use distillation, as if Anthropic or OpenAI are above that or something. Not to mention that they basically ignored copyrights on all the data the siphoned and are now crying that people aren’t respecting their terms of use.
The reality is that China tops the world in artificial intelligence publications today. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. Anybody who thinks China is simply distilling glorius American models is not engaging with reality.
Not to mention that US companies models constantly distill each other as Musk was forced to admit under oath. This whole narrative has just been a massive cope.
China also has a centralized data repo that companies can access but must also contribute to
Yup, and a massive advantage in the amount of data available by virtue of having a much bigger population.