← Back

Google delays flagship AI while spamming cheap Gemini models

Original version ·

Rather than delivering the promised super-brain, Google decided to flood the market with budget-friendly alternative models while quietly pushing back its primary powerhouse.

The tech giant pivoted away from a massive flagship release to focus on cutting costs and boosting processing speed across its lineup. The new Gemini 3.6 Flash was designed specifically to handle code generation, agent workflows, and office tasks while drastically cutting down on wordiness. Tests show it cuts output token usage by 17% in general workflows and up to 65% in coding benchmarks, pricing out at $1.50 per million input tokens and $7.50 per million output tokens.

To serve hyper-budget applications, Google also launched Gemini 3.5 Flash-Lite, which processes queries at a staggering 350 tokens per second for just $0.30 per million input tokens. It trades deep reasoning capabilities for raw speed, targeting massive customer-service chatbots and basic data extraction pipelines that usually burn through cash like standard corporate subscriptions.

The third addition, Gemini 3.5 Flash Cyber, focuses exclusively on hunting software bugs and feeds directly into Google's internal agent CodeMender. Instead of a public rollout, access is strictly limited to government entities and vetted partners. During tests on the V8 engine, it detected 55 unique security bugs, outperforming standard models like Claude Opus 4.6 which only caught 36 vulnerabilities.

Meanwhile, the ambitious Gemini 3.5 Pro hit a wall, as reports revealed Google postponed its public release after disappointing coding test results. However, engineers have already kicked off training for Gemini 4, claiming it represents the single largest training run in the history of the company.

When the flagship brain stumbles over basic code, spitting out three faster versions of the same thing seems like the ultimate tech magic trick. The tech landscape continues to trade quality breakthroughs for volume, proving that if a system cannot solve the problem, doubling down on cheaper workers is always an option.

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

0/24
  1. No comments yet.