Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters

Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters Chinese AI startup Z.ai (formerly Zhipu AI) has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. The model was previously known in stealth as "Ox Alpha" and topped OpenRouter's leaderboard before its official release. GLM-5.3-Flash features a hybrid architecture combining sparse and linear attention with Manifold-Constrained Hyper-Connections (mHC), redu

2 min
Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters

Z.ai releases GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters

Chinese AI startup Z.ai (formerly Zhipu AI) has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. The model was previously known in stealth as "Ox Alpha" and topped OpenRouter's leaderboard before its official release.

Cover image: Z.ai GLM-5.3-Flash, a 320B parameter hybrid sparse-linear attention model with 18B active parameters, native multimodal (text, image, video), first multimodal GLM-5 series model

GLM-5.3-Flash features a hybrid architecture combining sparse and linear attention with Manifold-Constrained Hyper-Connections (mHC), reducing long-context serving costs while preserving precise long-context capabilities. With 320B total parameters and just 18B active parameters, it delivers improved efficiency. The model was pre-trained on a 30T-token multimodal corpus and supports text, image, and video input.

The company claims GLM-5.3-Flash outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. Pricing on the Z.ai API Platform is $0.15 per million input tokens and $0.50 per million output tokens.

The model weights are available for download on Hugging Face (zai-org/GLM-5.3-Flash) and can be deployed locally with vLLM, SGLang, TokenSpeed, and KTransformers frameworks.

Architecture illustration: GLM-5.3-Flash architecture, hybrid sparse-linear attention, mHC connections, multimodal (text and image), 320B total/18B active parameters

Sources:

  • Techmeme: Z.ai releases GLM-5.3-Flash, the first natively multimodal GLM-5 series model, with 320B parameters, saying it outperforms GLM-5.2 at "one-tenth the price" (Z.ai) https://www.techmeme.com/260826/p35#a260826p35
  • Bloomberg: China's Z. AI made Ox Alpha stealth model that rivals DeepSeek https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek
  • TechCrunch: Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model https://techcrunch.com/2026/08/26/surprise-z-ai-is-the-ai-lab-behind-the-mysterious-ox-alpha-model/
  • Hugging Face: zai-org/GLM-5.3-Flash model card https://huggingface.co/zai-org/GLM-5.3-Flash

Written by

More to read