OpenAI Just Nuked Their Own Favorite AI Benchmark: It's 30% Broken Trash
In a display of absolute peak corporate irony, OpenAI spent months pushing SWE-Bench Pro as the industry standard for AI coding skills, only to perform an audit proving it's fundamentally flawed. Watching the benchmark king shoot its own foot is comedy gold.
OpenAI performed a deep-dive audit into SWE-Bench Pro, a darling of the AI evaluation world, and discovered that roughly 30% of its coding tasks are essentially broken. After betting the house on this benchmark, they have officially retracted their endorsement, suggesting that model developers have been chasing shadows for months.
The methodology was simple but brutal: OpenAI used automated filters and teams of human engineers to vet the tasks. Whether it was the robots or the humans doing the checking, the result was consistent: over a quarter of the benchmark is technically invalid. The issues range from tests requiring non-existent specs to bizarre mismatches where hidden tests demand formatting that contradicts the prompt instructions.
A perfect example of this technical gaslighting involves an OpenLibrary serialization task. The prompt explicitly asks for a specific format with a single space, while the hidden test script demands two. Naturally, any AI smart enough to follow instructions gets penalized for being correct, while the benchmark marches on as if it actually measures intelligence.
The core problem, according to OpenAI, is the lazy industry habit of scraping real-world GitHub pull requests and calling them benchmarks. These tasks weren't designed for AI; they were messy, human-centric debates that make for terrible standardized testing. The company is now pleading with the community to stop using existing code as a shortcut and start manually crafting high-quality evaluation sets from scratch.
It is truly heartwarming to see the industry finally admit that its "objective" performance metrics are mostly just glorified guessing games. When the companies setting the rules realize they have been grading their own homework with a broken calculator, it raises the question of whether any of these current AI capability leaps are anything more than statistical noise.
Source: OpenAI
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.