Google Launches EmbeddingGemma 2: Compact 740M Parameter Multimodal Embedding Model for Ultra-Low Latency Edge AI
Google 推出 EmbeddingGemma 2:超輕量 740M 參數,讓裝置端擁有強大「多模態語意搜尋」與即時決策力

Google DeepMind has launched EmbeddingGemma 2, a 740M parameter open-weight multimodal embedding model. Designed for privacy-first, on-device operations, it natively maps text, images, video, and audio into a single vector space. Running on as little as 567MB active RAM on a Pixel 11 Pro, it replaces complex model chains with a single, efficient engine. Developers can deploy it via LiteRT, MediaPipe, or ML Kit to enable real-time semantic search, visual retrieval, and <100ms on-device decision routing without fine-tuning.
Key points
Unified Multimodal Space
Natively maps text, images, video, and audio into a single vector space, eliminating complex pipeline chains.
Resource-Efficient Edge Footprint
740M parameters, running on ~191MB RAM for text-only and ~567MB for full multimodal capabilities on a Pixel 11 Pro.
Sub-100ms Zero-Shot Decision
Works as an instant on-device decision engine for intent routing, evaluating 500 options in under 100ms without fine-tuning.
Unified Developer Tooling
Supports LiteRT, MediaPipe, and Android ML Kit, clocking 37.3ms for visual embeddings on a MacBook M5 Pro GPU.
How it works
Why it matters
EmbeddingGemma 2 removes the high latency and privacy risks of cloud-based AI. Historically, local multimodal retrieval required chaining fragmented models (ASR, image captioning, and embeddings), which bloated local memory. This single, cohesive model enables instant offline media search and instant intent-routing, offering a highly practical blueprint for private, latency-sensitive edge processing on consumer hardware.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
- Startup Founder
How to use it
- 1On-Device Semantic Media Search
- 2Offline Meeting & Document Companion
- 3Real-Time Action & Intent Routing
Limitations & caveats
- Hardware Constraints: The 567MB active RAM requirement for the full multimodal model may still strain older or non-flagship edge devices lacking NPU acceleration.
- Limits of Zero-Shot Reasoning: While ultra-fast for simple intent routing, its lightweight architecture may struggle with highly complex reasoning tasks compared to cloud-based LLMs.
Related

Falcon-Emirati-7B: Bridging the Gap in Emirati Arabic Dialect and Culture
解鎖阿聯酋方言與文化:專為在地語境打造的 Falcon-Emirati-7B 模型
Falcon-Emirati-7B is a 7B parameter model specialized in Emirati Arabic, capturing local dialect, Nabati poetry, and cultural nuances where generic models fail.
Base Models Can Reason: Unlocking Latent Performance with Strategic Starting Tokens
基礎模型也能推理:啟動關鍵「開頭 token」釋放隱藏實力
A new study reveals that forcing base models to start with specific token cues like 'Okay' triggers reasoning behavior comparable to RL-tuned models, tracing this effect directly to pre-training data structures.
Towards Looped Models Done Right: Rethinking at Fixed Points for Efficient Training, Decoding, and RL
循環語言模型的「定點」重塑:邁向高效訓練、解碼與強化學習的全新架構
This paper optimizes looped language models near their fixed points using learned depth priors and orthogonal input injection, achieving up to 1.79x faster prefill, 2x faster RL training, and 3x smaller KV cache.