Z.ai Launches GLM-5.3-Flash With Vision, Video and a 1M-Token Window
Z.ai’s new open-weight model brings native multimodal input and agent-focused performance to a low-cost API tier.

Z.ai has released GLM-5.3-Flash, a multimodal model aimed at developers building coding agents, computer-use workflows and long-context applications. LLM Stats reports that the model launched on August 26, 2026, after initially appearing under the ox-alpha name on OpenCode and OpenRouter.
This is not simply a discounted version of the text-only GLM-5.3. Flash uses a 320B-parameter mixture-of-experts design with 18B active parameters, supports text, images and video natively, and offers a 1M-token context window. Its weights are available under the MIT license, with support for local serving through tools including vLLM and SGLang.
Why builders may care
The API list price is $0.15 per 1M input tokens, $0.03 for cached input, and $0.50 for output. A 50% introductory discount runs through September 9, 2026, at 16:00 UTC, bringing those rates to $0.075, $0.015 and $0.25. Developers should budget against the permanent list price rather than the temporary promotion.
Z.ai’s launch figures suggest meaningful gains for agent workloads: DeepSWE reached 63.4, versus 46.2 for GLM-5.2, while AutomationBench scored 48.8, up from 26.2. Terminal Bench 2.1 was closer, at 84.3 versus 81.0. The results are vendor-reported and have not been independently verified by LLM Stats.
For multimodal tasks, the model posted 78.0% on Chartography and 77.8% on MVBench, though its 53.4% BabyVision score trailed Gemini 3.7 Flash’s 70.9%. With always-on reasoning, open weights and low serving costs, Flash gives teams another option for visual agents and coding systems—provided they validate performance on their own workloads.
Source: LLM Stats
Comments
Log in to join the discussion