AI thinks like a biologist: OpenAI's test shows a massive leap
We might actually be getting closer to AI scientists that don't just copy-paste code. OpenAI just dropped a biology benchmark that tests actual research intuition, and the results are both wildly impressive and hilariously human.
The newly released GeneBench-Pro by OpenAI puts AI models through 129 synthetic tasks across 10 biological fields like pharmacogenomics and oncology. Instead of testing if an AI can run a pre-packaged script, this test evaluates if the system can distinguish actual biological patterns from noisy data junk.
To keep things honest, 82 of these tasks were heavily vetted by actual postdocs and professors to ensure they resemble messy, real-world science. The latest GPT-5.6 Sol model absolutely crushed the previous generation, skyrocketing from under 5% accuracy on the first GeneBench version to 28.7% on the maximum reasoning level.
This leap isn't just about guessing better; the model actually adjusted its mathematical approach on the fly, switching to a complex marginal structural model to handle drug-treatment feedback loops. Meanwhile, competitors are still struggling in the remedial class, with Claude Opus 4.8 reaching 16%, Gemini 3.5 Flash at 8.1%, and Grok 4.3 barely scratching 1.5%.
However, even the smartest model still acts like a distracted undergrad. The creators noted that when GPT-5.6 Sol spots a red flag in the data—like a technical glitch—it simply ignores its own discovery and keeps pushing forward with the original, flawed plan.
According to Alexander Stradwick Young, an associate professor at UCLA, these tasks would give actual PhD students a headache. OpenAI did admit that they used their own GPT models to help design the test, which is a bit like a teacher letting their favorite student help write the final exam, but the massive performance gap suggests there's real brainpower under the hood.
It seems the dream of AI curing diseases while humans sit back and sip margaritas is still a few logic updates away. If the smartest silicon mind can notice an error and just choose to ignore it, we haven't built a super-intelligence—we've just successfully automated academic procrastination.
Source: OpenAI
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.