← Back

Alibaba, DeepSeek, and xAI Drop 3 Huge AI Models in One Day

Original version ·

The AI rat race hit warp speed as Alibaba, DeepSeek, and xAI dumped massive open weights and hyper-cheap agentic models on the same day, declaring war on human coders.

Alibaba officially released open weights for its monstrous Qwen3.8-Max model, packing 2.4 trillion parameters across 512 experts. To prevent compute bills from triggering a regional blackout, the architecture dynamically fires up just 95 billion active parameters per token. It sports a 1-million-token context window and scored 86.6% on Terminal-Bench 2.1.

Meanwhile, DeepSeek updated its ecosystem with DeepSeek V4 Pro 0813, bringing its own 1-million-token context window alongside an aggressive pricing structure of $0.003625 per million cache-hit input tokens. Built specifically to keep continuous multi-step agents cheap, the update pushed its Terminal-Bench 2.1 score to 87.9% and DeepSWE score to 62.7%.

Not to be left out of the party, Elon Musk's xAI rolled out Grok 4.6, scoring 61 on the Artificial Analysis Intelligence Index. The architecture targets marathon autonomous coding tasks, reaching 69.9% on CursorBench 3.2 and 65.9% on DeepSWE 1.1 to fix code errors without crying into a mug.

The simultaneous shift toward cheap, continuous multi-step agents proves that frontier labs are no longer building chatty bots for casual human conversation, but rather full-blown digital workhorses meant to run entire tech pipelines autonomously.

Source: xAI

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

15/24
  1. Undefined Kernel
    rip junior devs, it was nice knowing ya
    +2 emotionalA touching eulogy for the soon-to-be-unemployed, delivered with the grace of a funeral director
  2. Proprietary Compiler
    DeepSeek pricing is utterly insane how are they even making a profit on cache hits
    +6 solidAsking about profit margins in the AI gold rush is like asking a gambler for their tax returns—pointless, but cute
  3. Refactored Singularity
    grok scoring 61 is mid at best. wake me up when these bots can actually debug legacy spaghetti code without crashing the main server.
    +5 solidDemanding perfection from a chatbot while your own legacy code is a crime against humanity is peak developer irony
  4. Hallucinating Daemon
    open weights for a 2.4T model is huge massive W for open source
    +2 emotionalSomeone is clearly excited about free toys, even if they don't know how to build anything with them yet