Thinking Machines released Inkling-Small two weeks after its debut model, Inkling. The new model has 276 billion total parameters and 12 billion active parameters per token, compared to Inkling's 975 billion total and 41 billion active. It scores 40 on the Artificial Analysis Intelligence Index, one point behind the flagship. Full weights are on Hugging Face under Apache 2.0. Discounted API pricing starts at $0.58 per million input tokens.

The benchmark split is where the article earns its read. Inkling-Small beats Inkling on SWE-bench Verified (80.2% vs 77.6%), Terminal Bench 2.1 (64.7% vs 63.8%), and several science and reasoning evals. It loses badly on factual tasks: 15.5% on tau-cubed Banking versus Inkling's 23.7%, and a negative AA Omniscience score. That asymmetry tells you exactly where to deploy it and where not to. The architecture behind it, a 42-layer sparse Mixture-of-Experts decoder routing each token to 6 of 256 experts plus 2 shared experts, explains how a 276B model runs on 12B active parameters at inference time.

Small is relative. The BF16 checkpoint needs 600 GB of aggregate GPU memory across 4x NVIDIA B300 or 8x NVIDIA H200 GPUs. A quantized NVFP4 version drops that to roughly 180 GB and can run on a single B300. No laptops, no MacBooks, no gaming rigs. The Apache 2.0 license, not a custom restricted variant, matters as much as the hardware specs for enterprise buyers who need fine-tuning rights and commercial freedom without revenue thresholds or branding obligations.

[READ ORIGINAL →]