DeepSeek surpassed Google to become the second-largest model provider by token volume on Vercel's AI Gateway in July 2026, processing a quarter of all gateway tokens versus Google's 11%. Google ran nearly 40% of gateway volume in April. DeepSeek V4 Flash alone ran more tokens than all of Google combined, nearly a fifth of total gateway volume and 70% more than any other single model. Moonshot's Kimi K3, released July 16, reached eighth by token volume on its last full day of the month, running twelve times the tokens per request of its predecessor and capturing 82% of all Kimi traffic within two weeks.
Open-weight models crossed 36% of token volume in July but the real break was on the spend side: open-weight share of gateway spend more than doubled to 8.6%, the highest recorded, driven almost entirely by Kimi K3 and Z.ai's GLM 5.2. These are the first open-weight models winning workloads historically owned by closed labs, priced at more than eleven times DeepSeek's rate per token. Meanwhile, overall gateway spend grew 37% while volume grew 59%, pushing the average price per token down 13.6%. That decline came entirely from routing choices: among teams running over 10 million tokens in both June and July, three in five changed at least a quarter of their model mix.
Anthropic remains the structural outlier. It held 65.1% of all gateway spend on 30% of token volume, with average token prices 4.4 times the rest of the market. Claude Fable 5 returned July 1 after a three-week export-control suspension and grew to 13.2% of total gateway spend. In coding agents, the largest use case by tokens, Anthropic collected more than 80% of spend. The full report details how model switching, image and video market share shifts, and the concentration of price competition in the lowest-revenue tier of the market are reshaping where enterprise inference dollars actually land.
[READ ORIGINAL →]