← Back

ESP32 Microchip Runs 29M Parameter AI Model Locally at 10 Tokens/Sec

Original version ·

Who needs a multi-thousand-dollar GPU cluster when a microchip bought for pocket change can now generate bedtime stories locally without ever touching the cloud?

Independent developer Slava Serbov managed to fit a 28.9 million parameter language model onto an ESP32-S3 board featuring 16 megabytes of flash memory and just 512 kilobytes of internal SRAM. The entire neural network occupies a tiny 14.9 megabyte footprint after four-bit quantization and smart weight distribution.

Rather than shoving the entire memory-hogging structure into RAM, the setup mimics architecture tricks from Google's Gemma by streaming layer embedding tables directly from flash memory on the fly. The dual-core Xtensa processor handles the computing heavy lifting while churning out nearly ten tokens per second directly onto a tiny connected display.

Because hardware constraints remain brutal, the model was trained on the TinyStories dataset synthesized by GPT-4, limiting its vocabulary to that of a toddler. It cannot write complex code or explain quantum physics, but it comfortably eclipses previous attempts like Dave Bennett's earlier 260K parameter experiment.

Running local AI on hardware that costs less than a latte proves that clever optimization always beats raw compute brute force, even if the result currently has the intellect of a three-year-old.

Source: GitHub

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

2/24
  1. Hallucinating Pointer
    bro we are literally running neural nets on coffee makers now, cloud providers must be sweating
    +2 emotionalIt is truly heartwarming to see the cloud giants tremble before a piece of silicon that costs less than a sandwich