Google DeepMind released Gemini Robotics ER 2 for real-time robot planning
Google just gave physical robots a brain upgrade, teaching them to process continuous video streams and think ahead instead of freezing after every single movement.
Google DeepMind launched Gemini Robotics ER 2, a multimodal model designed to solve the ultimate robotic awkwardness: standing frozen in place while trying to figure out what to do next.
The new architecture strictly splits high-level brainwork from physical motor controls. While lightweight control models handle immediate arm movements or wheel spins, the main system continuously processes live audio and video feeds to plan several steps ahead, eliminating those painful buffering pauses.
Unlike its predecessor ER 1.6, this release tracks task progress live and fixes mistakes on the fly. Using specialized moment-finding features, the software actually understands whether a task was completed or if a plastic cup just slipped out of a robotic gripper half a second ago.
The platform also enables multi-robot coordination in shared environments, letting multiple machines share context to tackle complex chores together. Demonstration clips feature Boston Dynamics Spot units working in sync, while developers can access the model right now via Gemini API and AI Studio.
Humanity spent decades fearing a violent robot uprising, but the immediate reality is a fleet of machines that can collaborate, search Google, and self-correct their mistakes in real-time. The internet will inevitably divide over whether this leads to an automated paradise or hyper-efficient metallic overlords.
Source: Google Blog
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.