Alibaba Drops Qwen3.8-Max: 4x Cheaper Than Claude Opus 5 with Open Weights
Chinese tech giant Alibaba just crashed the expensive AI party with a massive open-weight threat that makes proprietary AI labs sweat profusely over their margins.
The Qwen team at Alibaba deployed its new flagship 2.4-trillion parameter Mixture-of-Experts model, activating 95 billion parameters per token while maintaining a 1-million token context window. The infrastructure architecture builds directly upon the foundation laid by Qwen3.5.
API access was set at $2 per million input tokens and $6 per million output tokens, undercutting competitor rates while developers promised to publish the full model weights on Hugging Face and ModelScope next week alongside Qwen3.8-27B.
Independent evaluations on Terminal-Bench 2.1 showed Qwen3.8-Max scoring 86.6%, placing it slightly behind Kimi K3 at 88.3% and Claude Opus 5 at 89.1%. On the DeepSWE v1.1 programming benchmark, the gap widened with the model trailing Kimi K3 by 10.9 percentage points.
Testing on Humanity's Last Exam yielded a 56.2% score with tools enabled, placing performance on par with rival Chinese flagships despite the lower operating costs.
Demonstrating autonomous execution capabilities, an agent named oh-my-cli ran on GitHub for 16 days inside an empty TypeScript repository, generating 265 commits, 127 pull requests, and 151 issues without human supervision.
API integration supports OpenAI and Anthropic protocols for connection to Claude Code, Codex, and OpenClaw, featuring configurable reasoning depth in Qwen Studio and QwenCloud.
Proprietary AI vendors now face an uncomfortable choice between bleeding market share or sacrificing their lucrative profit margins. The relentless push toward cheap open weights continues to turn premium engineering intelligence into a commoditized utility.
Source: Qwen Blog
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.