Z.ai Drops GLM 5.3 Flash: An Open-Weights Beast That Makes Claude Sweat
While everyone was busy debating closed-garden models, Z.ai just nuked the status quo. GLM 5.3 Flash has landed with open weights and a price tag that feels like a clerical error. It is essentially a high-performance bargain bin special for your GPU.
The AI scene just got a bit more crowded with the arrival of GLM 5.3 Flash, the first native multimodal model in the GLM-5 lineup. Z.ai made the bold move of throwing the model weights onto Hugging Face under a generous MIT license, effectively inviting the entire internet to dissect their creation. Before this official coming-out party, the model was quietly terrorizing the leaderboards under the codename Ox Alpha, where it managed to dominate traffic on OpenRouter while wearing a digital disguise.
Under the hood, the architecture relies on a mixture-of-experts approach, packing a massive 320 billion parameters total but only firing up 18 billion per token to keep the lights on and the electricity bill manageable. The training process gobbled up a staggering 30 trillion tokens of text and visual data to reach its current state. On the technical front, it posted an impressive 84.3 on Terminal-Bench 2.1, 63.4 on DeepSWE, and 55.3 on HLE with tools.
Z.ai claims the model officially dusts its predecessor GLM 5.2 while breathing down the neck of Claude Opus 4.8 in coding tasks, all while operating at a fraction of the cost—specifically, a ten-fold reduction in API pricing. It has already been integrated into KodaCode for users of VS Code, JetBrains, and CLI interfaces. The audacity of releasing such a capable beast for free use is either a brilliant strategy to own the developer ecosystem or a desperate cry for attention in a market saturated with hype-hungry chatbots that cost a fortune to run.
Source: The Next Web
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.