Moonshot AI has released Kimi K2, a 1-trillion-parameter mixture-of-experts model with 32 billion parameters active per forward pass, trained on roughly 15.5 trillion tokens. It posts competitive scores against Anthropic's Claude Sonnet 4 and OpenAI's GPT-4.1 on agentic and coding benchmarks, and Moonshot is offering it via API at $0.15 per million input tokens, undercutting most Western competitors on price.
The architecture and the pricing are both worth examining closely. Kimi K2 uses a custom optimizer called Muon, not AdamW, and Moonshot claims this reduced training instability at scale. The model is open-weight under a non-commercial license, which means developers can inspect and run it, but not build commercial products without a separate agreement. That distinction matters as enterprises evaluate supply chain risk in their AI stacks.
The original piece breaks down the benchmark comparisons model by model and digs into what Muon actually changes at the optimization level. If you are tracking how Chinese labs are narrowing the gap with frontier Western models, and doing it cheaper, this is the data you need to read directly.
[READ ORIGINAL →]