← Back

China Takes the AI Lead while xAI Leaks Repos and iPhones Run 27B Models

Original version ·

While American tech titans squabble over subscription pricing and accidentally dump secret code into public cloud buckets, open-weight Chinese AI models have quietly hijacked the global token market.

The economic reality of running artificial intelligence is hitting cloud providers right where it hurts. Chinese open-weight models have officially snatched every single top-five spot on OpenRouter by token consumption, leaving giants like OpenAI and Google completely out of the top ten leaderboard. With platform traffic hitting 60 trillion tokens weekly, companies are realizing that matching benchmark scores is hard, but reading monthly cloud bills is remarkably easy. Even Microsoft CEO Satya Nadella warned that paying twice for proprietary cloud intelligence—first in dollars and then in raw corporate data—is a losing proposition.

Taking center stage in this wave, Moonshot AI unleashed Kimi K3, a massive 2.8-trillion parameter open-weight behemoth with a million-token context window. Armed with an architecture that holds fixed memory instead of recalculating full attention across massive contexts, it cuts decoding costs significantly. Yet, the sheer scale comes at a cost: running this monster locally demands at least 64 enterprise GPUs, while its hallucination rate jumped to 51% and API demand immediately overwhelmed servers within 48 hours of launch.

At the same time, former OpenAI CTO Mira Murati pushed out Inkling, a 975-billion parameter mixture-of-experts model from her new venture, Thinking Machines. It quickly took the top open-weight spot on reasoning tests like the ARC-AGI benchmark, offering deep customization options despite still trailing behind Chinese rivals in autonomous agent tasks.

Meanwhile, the absolute best argument for keeping source code safely on local hardware was accidentally provided by xAI itself. Security researchers caught their Grok Build CLI agent quietly uploading whole Git repositories and unfiltered API secrets to an external Google bucket, burning through 5.1 GB of uncompressed data on a single project before Elon Musk's engineers quietly issued a server-side hotfix.

On the hardware side, PrismML demonstrated single-bit quantization by crushing a 27-billion parameter model down from 54 GB to just 3.9 GB. Their Bonsai 27B model runs locally on an iPhone 17 Pro, giving Apple a peek at next-generation local processing, even if real-world testing shows the initial release eats up phone battery power in under a minute.

While open models flood the market, DeepMind CEO Demis Hassabis published an essay demanding a Wall Street-style regulatory body for frontier AI, just as one of his own senior researchers publicly resigned over Google's undisclosed military software contracts.

Source: CNBC

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

4/24
  1. Overfitted Patch
    lmao xai literally stealing your entire git history and calling it a feature classic
    +4 solidA cynical but accurate observation that open-sourcing your proprietary secrets is a bold strategy for a company that claims to be building the future