Mistral drops Le Chonk: A 1.05 Trillion Parameter Monster is Here
Mistral just birthed a trillion-parameter behemoth, because apparently, humanity hasn't reached its peak of digital gluttony yet. It is essentially a giant brain that claims to be great at everything, provided you have a spare continent to power it.
Le Chonk, formally known as Mistral Large 4, hits the scene as a multimodal MoE model packing a staggering 1.05 trillion parameters. Of that massive digital mountain, only 49 billion parameters are actually active during any single calculation, proving that even AI likes to be lazy when it can. The package also includes a 1.6-billion-parameter vision encoder to help the model pretend it can actually see the chaos it helps create.
The company trained this beast from scratch over two months using roughly 4,000 Nvidia Grace Blackwell GPUs housed in their own European data centers. On the DeepSWE v1.1 benchmark, the model posted a score of 62%, technically edging out GLM-5.3 and DeepSeek V4 Pro in internal tests. While the full weights aren't dropping until October 27, the API is open for business with a context window of up to 1 million tokens.
Pricing kicks off at $1.36 per million input tokens, though early adopters get a 50% discount for the first two weeks. It seems the race to the bottom of our wallets is just as competitive as the race to the top of the parameter count.
The shift toward these massive models suggests that intelligence is now measured strictly by how many billions of dollars in electricity can be burned through in a single training run. Whether this leads to actual innovation or just a very expensive way to generate hallucinations remains the industry's favorite guessing game.
Source: Mistral AI
Comments
Help shape the next version: Add context or suggest a correction. AI review can add points toward a rewrite. Reviews and updates may take time; a full meter does not guarantee a new version.