GLM-5.3-Flash Launch Details
Z.ai formally launched GLM-5.3-Flash, revealing it as a natively multimodal model with a 1M-token context window. The model utilizes 320B total parameters, with 18B active parameters, and is released under the MIT License. Z.ai made the model available via weights, API, chat, coding plan, and AutoClaw, offering a comprehensive deployment option.
Performance Benchmarks
Internal benchmarks within the Z.ai Code Bench indicate GLM-5.3-Flash outperforms GLM-5.2 at every effort level and is on par with Claude Opus 4.8 on coding tasks. Independent evaluations, conducted by Artificial Analysis, placed the model’s Intelligence Index score at 57, tying it with GPT-5.6 Terra and Muse Spark 1.2 while significantly reducing the cost per task. The API price is set at $0.15 / 1M input, $0.50 / 1M output, with cached input costing approximately $0.026–$0.03 / 1M.
Architectural Changes and Efficiency
Rasbt provided a detailed architectural breakdown, noting a shift from GLM-5.2’s 744B-A40B backbone to 320B-A18B. The model is described as "super hybrid" due to the use of efficient attention variants. The release also incorporates improvements in visual intelligence, potentially driven by coding/RL style training, and a system-oriented agent co-authored parts of the work.
Agentic Task Performance
Artificial Analysis reported that GLM-5.3-Flash performs competitively on agentic tasks, tying with GLM-5.3 and Grok 4.6 on Terminal-Bench v2.1 and achieving 47.2% on τ³-Banking, trailing GLM-5.3 by 3.1 percentage points. This suggests a strength in practical code/agentic workflows compared to broad factual knowledge.
Source: https://www.latent.space/p/ainews-nvidia-buys-huggingface-for



