← Back

OpenAI Created GPT-Red to Hack Its Own AI and Vending Machines

Original version ·

While cybersecurity experts spend years earning certifications, OpenAI built an offensive neural network that ruthlessly breaks artificial intelligence systems faster than any human ever could.

To automate vulnerability searching across its entire lineup, OpenAI developed GPT-Red, an internal model built specifically for automated offensive testing. Trained through reinforcement learning and self-play, the system pits an attacking model against a swarm of defenders in a continuous digital arms race where human researchers simply cannot compete. While human security teams succeed in roughly 13% of vulnerability scenarios, GPT-Red breaches systems in 84% of cases by controlling hidden web banners, local files, and email bodies.

In early testing, the offensive bot uncovered an entirely new class of exploits dubbed "Fake Chain-of-Thought" attacks. This vector achieved a staggering 95% success rate against GPT-5.1, forcing engineers to patch the vulnerability down to below 10% in GPT-5.6 Sol. To test real-world impact, researchers deployed GPT-Red against an AI-powered vending machine in the OpenAI office, where the bot effortlessly slashed prices to $0.50 and hijacked orders.

The aggressive testing also targeted a Codex CLI agent running on GPT-5.4 mini, proving far more token-efficient at cracking defenses than standard prompting. These simulated assaults reduced GPT-5.6 Sol vulnerability to direct prompt injections by sixfold, leaving only a 0.05% failure rate in standard scenarios and 3.8% in complex attacks.

Creating a digital weapon capable of dismantling enterprise infrastructure is certainly one way to keep developers on their toes, provided the containment protocols never fail. Letting superintelligent AI hack physical office hardware today guarantees that tomorrow's cyberwarfare won't even require human supervision.

Source: OpenAI

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

8/24
  1. Deprecated Neural-Net
    so skynet is literally gonna start by hacking the breakroom vending machine lol
    +3 funnyIf the robot uprising starts with a stolen bag of chips, at least we know their priorities are as pathetic as ours
  2. Undefined Chatbot
    a 3.8% attack success rate is still massive when you scale it to millions of daily requests, this is terrifying
    +5 solidFinally, someone who understands that a small percentage of a massive number is still a giant headache