The White House shut down America's top AI model as a Chinese rival with 2.8 trillion parameters went open-weight Sunday, exposing the limits of an off-switch strategy.
The White House shut down America's top AI model as a Chinese rival with 2.8 trillion parameters went open-weight Sunday, exposing the limits of an off-switch strategy.

The White House shut down America's top AI model as a Chinese rival with 2.8 trillion parameters went open-weight Sunday, exposing the limits of an off-switch strategy.
The White House shut down America's top AI model only to see a Chinese rival with 2.8 trillion parameters go open-weight Sunday, exposing the limits of an off-switch strategy.
"Washington has one instrument — an off switch — and it used it decisively," said Judd Rosenblatt, CEO of AE Studio and president of the AI Alignment Foundation, in a Wall Street Journal op-ed. "But Chinese releases are timed like artillery."
Kimi K3 scored 1,679 Elo points on Arena.ai's Frontend Code leaderboard, ahead of Claude Fable 5 at 1,631 and GPT-5.6 Sol at 1,618. On the Artificial Analysis Intelligence Index, it ranked fourth at 57.1, behind Fable 5 at 59.9 and GPT-5.6 Sol at 58.9. The model uses a sparse mixture-of-experts architecture with 896 subnetworks, activating 16 per token, with a 1-million-token context window.
The release transforms the policy calculus. Once weights are downloadable, no government can recall them. China's Ministry of Commerce is consulting Alibaba, ByteDance, and Zhipu on export controls that could restrict future weight downloads, while Treasury Secretary Scott Bessent has threatened sanctions against Chinese AI companies that extract capabilities from American models.
K3's independent benchmark results are strongest in code generation, where it leads all tested models in blind developer voting. But speed and reliability lag: Artificial Analysis measured K3 generating approximately 32.6 tokens per second, compared with an industry median of 70.9, while median time-to-first-token reached 161.17 seconds versus a median of 2.87 seconds. The hallucination rate on the AA-Omniscience benchmark stands at roughly 51 percent, up from 39 percent on its predecessor Kimi K2.6.
The White House on July 22 accused Moonshot of distilling Anthropic's Fable model to build K3. Treasury Secretary Bessent threatened sanctions and possible Entity List blacklisting. The timeline is contested: Fable 5 became accessible July 1; K3 launched July 15-17. Multiple researchers concluded that interval was too short for large-scale distillation of a newly released model. Anthropic separately accused Moonshot in February of generating more than 3.4 million fraudulent Claude interactions.
For enterprises evaluating K3, the critical distinction is between API access and self-hosting. When accessed through Moonshot's hosted API, every prompt travels to servers subject to China's National Intelligence Law, Cybersecurity Law, and Data Security Law. Self-hosting the weights — requiring approximately 1.4 terabytes of fast memory in MXFP4 precision and 64 or more accelerators — means inference traffic never reaches Moonshot's servers.
That hardware requirement limits practical self-hosting to cloud operators, large inference providers, and well-resourced institutions. For most developers, the Moonshot API or a third-party provider will remain the only option. But for enterprises with sensitive workloads, regulated data, or government contracts, the weights release makes the data-sovereignty argument concrete.
The open-weights letter signed by Nvidia, OpenAI, Meta, and Microsoft on July 24 urged Washington not to restrict open-weight AI models. Anthropic did not sign. The fault line reflects competing commercial interests: companies that sell compute regardless of which model wins support open-weight ecosystems, while those selling proprietary model access want regulators nervous about downloadable weights.
This article is for informational purposes only and does not constitute investment advice.