Stories about GLM-5.3-Flash
2 related stories
GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia
AI InsightZ.ai has released the open-source model GLM-5.3-Flash, which has 32 billion parameters and scores only three points lower than the larger GLM-5.3 on the Artificial Analysis Intelligence Index, but at one-seventh the cost. Notably, all inference traffic runs on Chinese AI chips rather than Nvidia hardware. The release of GLM-5.3-Flash marks a new breakthrough in the balance between cost and performance of AI models.Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
AI InsightZ.ai has introduced GLM-5.3-Flash, the first native multimodal model in the GLM-5 series. It has 320 billion total parameters and 18 billion active parameters, supporting a context window of 1,048,576 tokens. The model weights are released under the MIT license on Hugging Face, with API pricing at $0.15 per million inputs and $0.50 per million outputs. GLM-5.3-Flash scored 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1.