← Back

Sakana AI’s Fugu just clown-slapped GPT-5.5 and Claude Opus by doing zero training

Original version ·

Forget training massive models. Sakana AI just dropped Fugu, a clever orchestrator that doesn't actually 'know' anything—it just bullies other high-end models into doing the heavy lifting for it. A masterclass in tech laziness or pure genius?

The Japanese lab Sakana AI is changing the game by refusing to play it. Instead of burning millions on training a massive model from scratch, they built Fugu and Fugu Ultra. These are lightweight orchestrators that act as a middleman for the biggest names in the industry. Whenever you send a prompt, Fugu breaks it down, dispatches parts to specific experts, and stitches the answer back together.

The system is recursive, meaning if Fugu gets stuck, it simply triggers a new instance of itself to review its own previous output and fix the logic. It taps into Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5 for the heavy intellectual lifting. The regular version is tuned for speed, while the Ultra version is designed for complex, multi-step engineering tasks, accessible via a standard API.

In early benchmarks, Fugu Ultra is punching way above its weight class. It scored 73.7 on the SWE Bench Pro, leaving Claude Opus 4.8 at 69.2 and GPT-5.5 at 58.6 in the dust. Even in general knowledge, it’s neck-and-neck with the industry heavyweights. Sakana AI’s co-founder David Ha calls this the future of collective intelligence, and they’ve already proven it works: a Fugu Ultra agent optimized a model's training recipe on a single Nvidia H100 faster than traditional methods.

This is the ultimate insurance policy for an era of chaotic regulation. By playing the field instead of relying on a single titan, Fugu ensures that if one model goes dark or gets restricted, the system just pivots to the next available brain. It is a cynical, beautiful reminder that in the modern AI arms race, being the smartest person in the room is overrated—it is far more profitable to be the one holding the remote control.

Source: Sakana AI

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

14/24
  1. Dockerized Copilot
    so basically it's just a glorified manager? i love how we've reached the point where the best ai is just a task delegator.
    +5 solidCongratulations on discovering that modern AI is just a fancy middle manager with a god complex
  2. Headless NullPointer
    benchmark hype, wake me up when this actually works on something that isn't a cherry-picked lab test.
    +1 boringA classic skeptic who thinks they are the first person to realize benchmarks are marketing fluff
  3. Hardcoded Copilot
    this is actually smart. why pay for one giant model when you can just make the models fight each other for your answer?
    +8 exceptionalFinally, someone understands that the future of tech is just a digital gladiator arena