Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
EmbeddingGemma 2 brings multimodal search to the edge - OpenSmartRoute
EmbeddingGemma 2 brings multimodal search to the edge
Google launched EmbeddingGemma 2, a compact open-weight model that maps text, images, and audio into one vector space. It runs locally on phones with minimal RAM and enables instant on-device semantic search.
Key points
EmbeddingGemma 2 fits in a 740M parameter footprint.
Text-only weights run on ~191MB of active RAM.
The model matches inputs to labels without training data.
Vision embedding latency is as low as 37.3 ms on Mac.
Why it matters: Developers can now build private, offline search and decision systems that run instantly on user devices without cloud costs.
By OpenSmartRoute editorial · written through the router by writer-small
From Google AI developers blog - “Bring multimodal semantic search to the edge with EmbeddingGemma 2”
Google DeepMind launched EmbeddingGemma 2 for edge devices today. This model maps text, images, and audio into one vector space. It runs locally on phones with minimal RAM usage. You can now do instant on-device semantic search without internet.
The model is compact and designed for local applications. It reduces the need to chain separate image captioning models. Speech-to-text and text-embedding models no longer need to work together. This cuts down latency and memory overhead significantly. Developers building private search tools will find this very useful.
Announcement - Google launches EmbeddingGemma 2 for edge devices
Google DeepMind released a new model called EmbeddingGemma 2. It is an open-weight multimodal embedding model. The team says it is best-in-class for its size. You can run it on local devices for privacy-first applications. This launch happens on October 6, 2026, according to the blog post.
The primary goal is to bring multimodal search to the edge. Edge means processing data directly on your phone or laptop. This avoids sending sensitive information to a cloud server. It also speeds up responses by removing network delays. The model acts as an ultra-low-latency decision engine.
You do not need training data to use it effectively. It matches user inputs against classification labels instantly. This feature is called zero-shot intent routing. It works in a matter of milliseconds. Google plans to make this available as a service soon. Android users will get access through ML Kit later.
Model Specs - Size, RAM usage, and multimodal capabilities
EmbeddingGemma 2 has a footprint of 740 million parameters. This is very small compared to large language models. It covers text, vision, and audio modalities in one package. The model uses modular encoders that you can load selectively.
Text-only weights require about 191MB of active RAM. The full multimodal model needs roughly 567MB on a Pixel 11 Pro. This fits easily into the memory of modern smartphones. Mid-range consumer hardware can now run this vector search. You do not need expensive enterprise-grade servers for basic tasks.
The model compresses weights using Quantization-Aware Training (QAT). It stores integers instead of floating-point numbers. This brings multimodal vector search within reach of standard devices. INT4 and INT8 quantization formats are the standard here. These formats reduce storage needs while keeping accuracy high.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
Reflection released Beam, a 501 billion parameter text-only model for coding and science. Apache 2.0 weights are available this month after training on 23.8 trillion tokens.
Zero-Shot Routing - How the model works without fine-tuning
Zero-shot routing means the model understands intent without extra training. It matches inputs directly against stored classification labels. This allows it to act as a decision engine on your device. You can route tasks based on what the user is asking for.
The system delivers instant results in milliseconds. It eliminates the need for fine-tuning specific datasets. This makes deployment faster and cheaper for developers. You get a ready-to-use model from Google DeepMind. No complex data preparation or training pipelines are required.
This capability replaces the need for multiple specialized models. Instead of separate tools for text, image, and audio analysis, you have one unified space. The model maps all these inputs into a single vector representation. This simplifies the architecture of any local search application you build.
Google AI Edge Gallery - Instant Media Search demo details
Google AI Edge Gallery is an interactive showcase app for mobile devices. It lets users test on-device AI models in practical settings. Two new demos powered by EmbeddingGemma 2 are now available. These include Instant Media Search and Video Moments Finder. You can download the latest version from the Google Play Store or Apple App Store.
Instant Media Search finds specific images or videos using natural language. It converts both your query and the media into embedding vectors. The system stores these vectors in a local SQLite database. It retrieves content by returning items with the largest cosine similarity. This creates an instant, interactive search experience as you type.
Search-as-you-type updates results live on every keystroke. If you type "Katze", the grid shows all feline photos immediately. Continuing to type "Katze schläft auf Tastatur" re-ranks the results in real time. This dynamic behavior brings the specific photo you want to the top. You can also select an image from your phone or use your live camera stream.
Video Moments Finder - Keyframe indexing and natural language search
Video Moments Finder helps you locate specific visual moments in local recordings. It does not require transcribing audio or generating intermediate text captions. This saves time and keeps processing entirely on the device. You select a local video file, such as a family video or sports reel.
The engine indexes vision with audio chunks to create embeddings for them. When you enter descriptive queries like "kids laughing", it computes a prompt embedding vector. It compares this against the video frame embedding instantly. The system highlights exact timestamps where matching moments occur.
This feature works without needing to watch the entire video first. You can search for events like "dog catching a frisbee" or "person blowing out birthday candles". The natural language query drives the search process directly. The results appear as precise time markers on your video player.
Google AI Edge Foresight - Meeting companion features on Mac
Google AI Edge Foresight is an experimental app for Mac computers. It acts as a context-aware meeting companion directly on your device. It assists with note-taking, indexing, and recalling conversation transcripts. You can use it completely offline without any network connection.
Foresight integrates directly with your system audio and microphone. This allows it to work out-of-the-box with any meeting platform. It processes sensitive information locally so data never leaves your machine. The model runs on fully local AI processing power.
Enhanced Note-taking enriches manual shorthand notes in real-time. Cross-Modal Retrieval lets you find different types of media using natural language. Your private files and conversation transcripts remain secure throughout the meeting. There are no cloud subscription costs associated with this feature.
ML Kit Integration - Coming support for Android apps
Google plans to make EmbeddingGemma 2 available as a service on Android soon. This will happen through ML Kit, which provides standardized platform integration. Developers will get production-grade quality without managing the model lifecycle themselves. APK bloat is eliminated by this unified approach.
ML Kit features blazing fast performance when an NPU is available. Automatic model updates ensure you always have the latest version. The integration simplifies deployment for Android app developers significantly. You do not need to worry about complex model hosting or maintenance.
This support targets a broad range of edge devices. It optimizes performance across various hardware configurations. The goal is to provide simplicity without sacrificing speed. Developers can focus on their application logic rather than infrastructure.
MediaPipe Tasks - Unified interface for cross-platform development
MediaPipe Tasks adds support for EmbeddingGemma 2 to existing tasks. This includes the Embedder and Semantic Retriever components. It creates a unified interface that abstracts low-level data transformations. Developers can write code once and deploy across all platforms.
The MediaPipe Universal Embedder Task handles image resizing automatically. It performs tensor normalization and multimodal tokenization for you. You pass raw images or text strings to receive normalized vectors. The output is 768-dimensional or truncated versions like 128d or 512d.
The MediaPipe Semantic Retriever Task runs fast Approximate Nearest Neighbour (ANN) searches. It indexes embeddings on-device and returns ranked matches in single-digit milliseconds. This speeds up retrieval compared to traditional search methods. The engine is optimized for edge environments specifically.
EmbeddingGemma 2 also works with the MediaPipe Decision Task. This allows real-time classification of image, text, or audio inputs. You can make instant decisions on-device without fine-tuning. Use cases include routing actions or predicting user behavior. A chess game example showed less than 100ms response time for evaluating options.
Performance Benchmarks - Latency numbers across CPU, GPU, and NPU
Visual embeddings take as little as 37.3 ms on a MacBook M5 Pro GPU. This translates to processing 26.9 images per second at that speed. The benchmark evaluated with a maximum budget of 70 vision tokens per image. These numbers show the efficiency of the model on high-end hardware.
Latency varies across CPU, GPU, and NPU backends depending on the device. LiteRT delivers robust performance across a broad spectrum of edge devices. It supports web assembly (WASM) for browser-based inference as well. The engine simplifies deployment by using a single .litertlm file format.
Targeted optimizations exist for compute-intensive workloads like vision encoding. These optimizations unlock highly responsive, real-time AI experiences. Search-as-you-type functionality benefits greatly from these improvements. Visual embeddings remain fast even on variable text length inputs.
Why it matters
This technology lowers the cost of building private search applications. It removes the need for expensive cloud infrastructure and data transfer fees. Speed improves because processing happens locally without network delays. Safety increases since sensitive user data never leaves the device. Developers gain a powerful tool for creating responsive, offline-first apps.
How OpenSmartRoute helps
A team routing requests through OpenSmartRoute gains immediate access to EmbeddingGemma 2 without buying a new license. The router treats this model as one entry in its catalogue alongside other providers and tools. It scores every candidate on quality, cost, speed and safety before sending traffic.
The system uses the specific RAM footprint of 191MB for text-only weights to optimize local deployment. EmbeddingGemma 2 fits into a 740M parameter footprint, allowing it to run on devices with minimal resources. Vision embedding latency stays as low as 37.3 ms on Mac hardware.
Requests containing personal data stay on an on-premises model due to hard safety rules. The input guard spots prompt injection before any request leaves the secure environment. A savings ledger shows exactly how much each routed request cost compared to the most expensive option.
What to do
Download the Google AI Edge Gallery app from your app store today. Try the Instant Media Search and Video Moments Finder demos immediately. Check the GitHub repository to learn how these features are implemented. Start building your own experiences using EmbeddingGemma 2 on your device. Monitor upcoming ML Kit releases for Android integration support.