← Back

New Claude Fable 5 replaces human freelancers in 16% of real tasks

Original version ·

Forget useless academic benchmarks that test if AI can pass high school history. We finally have a test that throws LLMs into the wild west of real freelance gigs, and the results are both terrifyingly impressive and deeply hilarious.

The Center for AI Safety (CAIS) and Scale Labs just dropped an update to their Remote Labor Index. Instead of giving AI models multiple-choice questions, they threw actual paid freelance jobs at them, covering everything from 3D modeling and web apps to data analysis and audio editing.

A human expert then compared the AI's output directly with what a real, paid human freelancer delivered. If the AI's work was at least as good as the human's, it got accepted. Under these brutal conditions, the brand-new Claude Fable 5 managed to successfully replace human freelancers in 16.1% of real-world orders.

This performance absolutely demolished the competition. The previous-gen Opus 4.8 only managed an 8.3% success rate, while GPT-5.5 trailed behind at a measly 6.3%. When this index started less than a year ago, the absolute best AI model could only handle 2.5% of tasks, meaning we are looking at a massive six-fold increase in actual, practical capability in just eight months.

The test did hit a bizarre geopolitical roadblock when the US government restricted access to Claude Fable 5 due to export controls. CAIS only managed to run 218 out of the 240 benchmark projects before their keys were revoked. Even if the model completely failed the remaining 22 tasks, its worst-case score would still sit at 14.6%, which still stomps every other AI on the planet.

Lest we get too excited about the robot uprising, the benchmark organizers noted that "accepted" doesn't mean flawless masterpiece. The bar was simply "not worse than what a real human sent to a client," which highlights how often human freelancers themselves deliver slightly questionable work.

For example, in a task requiring a 3D model of a ring with a modified gem cut, Claude Fable 5 easily outclassed older models, but still left the metal prongs looking like a middle-school art project.

The line between theoretical AI hype and actual economic disruption is officially evaporating. As algorithms quietly learn to match the messy, good-enough standard of the global gig economy, the real question is how long human workers can survive on "slightly less sloppy" as a competitive advantage.

Source: Center for AI Safety

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

2/24
  1. Sandboxed Repo
    great, now i have to compete with a bot that doesn't sleep or complain about 'mental health days' for a $15 logo gig.
    +2 emotionalIt is truly heartwarming to watch the gig economy turn into a digital gladiator arena where the robots always win