GLM-5.3, the frontier open source model from Chinese startup z.ai, is now live on the API at $1.40 per million input tokens and $4.40 per million output tokens: the same rate as its predecessor, GLM-5.2. Cached input runs $0.26 per million tokens, with cached-input storage listed as free for a limited time. The model debuted last week after reportedly detecting a previously unknown vulnerability in Cursor, signaling serious coding and agent capability.
At $5.80 combined per million tokens, GLM-5.3 sits below Grok 4.6 at $8.00, GPT-5.6 Terra at $14.00, GPT-5.4 at $17.50, Kimi K3 at $18.00, and Claude Opus 5 at $30.00. It is more expensive than DeepSeek-V4-Flash off-peak at $0.88 and GPT-5.6 Luna at $1.40, but z.ai is positioning GLM-5.3 as a frontier-class model, not a budget one. The API uses OpenAI Chat Completions-compatible protocol, with existing GLM Coding Plan subscribers currently limited to that interface.
Two things remain unresolved: the exact date model weights will be released publicly, and the licensing terms that will govern them. Those two details will determine whether GLM-5.3 becomes a serious open-weight competitor or stays a hosted-API story. The full article includes a 30-model pricing comparison table and the context around z.ai's capability claims, which is worth examining before the weights drop and the conversation shifts.
[READ ORIGINAL →]