Poolside released Laguna S 2.1 on Tuesday: a 118-billion-parameter Mixture-of-Experts model that activates only 8 billion parameters per token, scores 70.2% on Terminal-Bench 2.1, and beats DeepSeek-V4-Pro-Max (1.6 trillion parameters, 64.0%) and Nvidia Nemotron 3 Ultra (550 billion parameters, 56.4%). The weights are live on Hugging Face under the permissive OpenMDW-1.1 license. Pre-training started May 22 on 4,096 Nvidia H200 GPUs. Public launch followed in under nine weeks.
The release is a direct response to a gap Poolside names explicitly: no Western lab has released open weights in this size class since OpenAI's gpt-oss-120b last August. DeepSeek, Qwen, Kimi, GLM, MiniMax, and Tencent Hunyuan dominate Poolside's own comparison tables. Co-CEO Jason Warner states the problem plainly: the West needs open-weight models it can trust, run, and build on. Poolside's core business is deploying models inside government and defense security boundaries, where closed API access fails compliance and sovereignty requirements. Releasing competitive open weights is both ecosystem strategy and top-of-funnel for that high-security deployment business.
The piece is worth reading in full for two reasons. First, the token economics argument: long-horizon agents consume a mean of roughly 249,000 completion tokens per trajectory on Poolside's hardest benchmark, making inference cost a real budget variable, and Poolside is pricing the 1-million-context deployment at $0.10 per million input tokens and $0.20 per million output tokens on OpenRouter. Second, the transparency claim: Poolside published the complete, unedited trajectory of every trial in its final benchmark runs, every reasoning step, tool call, and shell command, a move with little precedent among major labs and a direct challenge to AI's growing benchmarking credibility crisis.
[READ ORIGINAL →]