Liquid AI Releases LFM2.5: A 2.6B Local Agent Running Offline on Phones
Sending private queries to expensive cloud servers might finally become optional as Liquid AI squeezes a multi-step autonomous assistant straight into consumer hardware.
Engineers at Liquid AI squeezed 2.6 billion parameters into their new model, LFM2.5-2.6B, training it on roughly 34 trillion tokens before passing knowledge down through on-policy distillation. Instead of relying on massive data centers to make simple choices, the architecture relies on specialized expert routines for math, coding, and tool invocation.
The training finished inside real agent environments like Hermes Agent and OpenClaw to ensure the model actually handles complex task chains without losing its digital mind. While large language models usually eat RAM for breakfast, this one maintains a 128K context window while staying under 2.5 GB of CPU memory usage on portable devices.
Hardware benchmarks show execution rates reaching 220 tokens per second on Apple's M5 Max and roughly 30 tokens per second on a standard smartphone. High-end hardware like Nvidia's H100 pushes throughput to almost 15,000 tokens per second under heavy parallel workloads.
Internal testing claims LFM2.5-2.6B outperforms larger rivals like Gemma and Qwen in instruction following and tool usage, even if massive models still rule complex coding tasks. Open-source integration arrived immediately across llama.cpp, MLX, vLLM, and weights are already live on Hugging Face.
Big tech spent billions convincing everyone that intelligence requires a subscription and a burning data center, yet pocket-sized silicon keeps quietly proving them wrong. Cloud API providers might want to start looking over their shoulders before local scripts replace their billing departments entirely.
Source: Liquid AI
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.