Google's EmbeddingGemma 2 is out: Is it actually 'Best-in-Class'?
Google DeepMind just dropped EmbeddingGemma 2, a shiny new multimodal toy meant to put everything into one data bucket. While the marketing department is busy popping champagne, the actual devs are realizing it might just be another case of over-promising.
The core concept is to cram text, images, video, and audio into a single 768-dimensional space, ditching the need for separate conversion layers. The model clocks in at 740 million parameters, but you can swap out pieces: the text component takes 270M, while vision and audio encoders add 170M and 300M respectively. It uses the Gemma 2 architecture with 24 layers and an 8192-token context window.
For edge devices, the footprint is impressively small. Google claims a memory usage of around 191 MB for text-only, jumping to 567 MB for the full multimodal setup on a Pixel 11 Pro. However, developers running it locally in GGUF stacks are seeing closer to 730 MB once runtime overhead hits. The model does show a solid bump in code-related benchmarks, jumping from 68.76 to 78.68 on MTEB Code, but multilingual performance remains largely stagnant.
Real-world testing reveals the cracks in the "best-in-class" narrative. Models like Qwen3-Embedding-0.6B are currently outperforming it in several standard benchmarks, proving that "small" doesn't automatically mean "better." Furthermore, if you try to run this in float16, it will likely return NaN errors or just degrade silently. You have to stick to bfloat16 or float32. Even the clever Matryoshka Representation Learning (MRL) tricks have limits; shrinking your vector size to 128 dimensions causes a massive hit to multimodal accuracy.
It seems that even with the massive resources of Google, you cannot simply wish away the laws of physics and information density. The industry loves slapping "SOTA" stickers on everything, but until independent benchmarks stop pointing at Qwen, this is just another tool that works fine if you don't expect miracles.
Comments
Help shape the next version: Add context or suggest a correction. AI review can add points toward a rewrite. Reviews and updates may take time; a full meter does not guarantee a new version.