Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Google Releases EmbeddingGemma 2 for Multimodal Search - OpenSmartRoute
Google Releases EmbeddingGemma 2 for Multimodal Search
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
Key points
EmbeddingGemma 2 is a sub-1B model under the Apache 2.0 license.
It scores 14% higher than EmbeddingGemma 1 on code retrieval benchmarks.
Matryoshka Representation Learning allows truncating vectors to 128 dimensions for six times storage reduction.
The model supports an 8,192-token context window across all input modalities.
Why it matters: Developers can now run a single lightweight model to index and search diverse data types locally without massive compute infrastructure.
By OpenSmartRoute editorial · written through the router by writer-small
Google released EmbeddingGemma 2 for multimodal search. This new open model handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types. The model is compact and runs under the Apache 2.0 license. Engineers can now deploy it locally without massive compute infrastructure.
Modern search applications need to work across diverse content types. These applications range from technical documentation to source code. They also include images, video clips, and audio recordings. Finding models that deliver strong retrieval accuracy is a major challenge. Developers often struggle with low latency across all these formats. Running such systems usually requires massive compute infrastructure. This new model aims to solve those specific problems efficiently.
EmbeddingGemma 2 provides a single, compact open model for developers. It delivers exceptional multimodal performance relative to its small size. The model is based on Gemma 4 technology. It maps text, code, images, video, and audio into a unified space. This space has exactly 768 dimensions for every input type. Its modular architecture lets users load only what they need at runtime. You can scale from 270 million parameters up to 740 million parameters.
The model uses modular encoders that project into a unified vector space. Each modality gets its own specialized encoder for processing. However, all inputs go through a shared backbone network. The resulting embeddings occupy the same dimensional space regardless of input type. This design replaces chained models with a more efficient single system. You do not need separate pipelines for different media types anymore.
Performance tests show significant improvements over the previous version. EmbeddingGemma 2 outperforms EmbeddingGemma 1 on code search tasks. It also excels at technical retrieval in complex environments. This makes it ideal for local codebase indexing projects. Agentic code search benefits greatly from this superior understanding. The gap in performance is described as significant by the authors.
Storage optimization relies on Matryoshka Representation Learning technology. This technique enables dynamic dimension truncation to save memory usage. You can truncate vectors from 768 dimensions down to 128 dimensions. This cuts vector database storage requirements while retaining much of the original quality. For example, text and code quality remains high at 256 dimensions. Image, video, and speech retrieval keeps about 95% quality at that level too.
Google launched EmbeddingGemma 2, a compact open-weight model that maps text, images, and audio into one vector space. It runs locally on phones with minimal RAM and enables instant on-device semantic search.
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
Why it matters
This capability reduces costs and complexity for local RAG systems handling mixed media. Engineers save money on cloud storage and compute resources. Managers see faster deployment times for new search features. Safety improves because models run locally without external dependencies. The unified space simplifies the integration of diverse data types.
Developers can install the model via the sentence-transformers library. They need version 6.1.0 or later to use this specific model. The installation command includes flags for image, audio, and video support. Users must also have the transformers library installed on their system. Running the full model requires about 740 million parameters in memory.
The code example shows how to load the full multimodal setup. It imports the SentenceTransformer class from the sentence-transformers package. The model loads all encoders for text, code, images, video, and audio. This configuration uses the maximum parameter count of 740 million. Users can omit unused modality encoders to minimize memory usage. They set vision_config or audio_config to None in config_kwargs.
To minimize memory usage, you only load what your data needs. Running text-only at 270 million parameters saves significant RAM. Adding vision brings the count to 440 million parameters total. Loading full multimodal pushes it to 740 million parameters. Disabled encoders are never loaded into memory during runtime. This applies savings to both the weights and peak allocation.
EmbeddingGemma 2 is trained with short task instructions to steer representations. These specific tasks include retrieval, classification, and generation. Developers set prompt_name in encode() to add custom instructions. For retrieval, queries and documents get different prompts applied. The similarity function then calculates the match between embeddings.
Step 3 covers embedding images, video, audio, and interleaved inputs. You pass media as a dictionary keyed by modality type. Text and media can be embedded together in one call. Mark where each item goes with special tags like <|image|>. The model handles cross-modal search effectively for these cases. One text query can match against a photo and sound recording simultaneously.
Interleaved inputs allow one embedding for a product listing with multiple media types. You mark the text, photo, and video sections clearly. The resulting embedding represents the entire mixed input accurately. Despite coming from different modalities, the embeddings occupy the same space. They can be compared on their semantic similarity directly.
Step 4 explains how to truncate dimensions using Matryoshka learning. Pass truncate_dim with values like 512, 256, or 128. Set normalize_embeddings=True to get shorter, unit-length vectors. Queries and documents must use the same dimension for comparison. You can set this at load time instead of per call.
Storing a million 768-dimensional vectors takes roughly 1.5 GB of memory. Truncating them to 128 dimensions requires just 250 MB of memory. That is a six times reduction in storage requirements. It allows you to store six times as many embeddings in the same budget. This makes it much easier to fit large indexes in memory or on-device.
All modalities share the 8,192-token context window at fixed rates. The maximums assume a single modality with no text input. Pass media as file paths for video files like MP4. Use URLs for images and audio files. In-memory PIL images, arrays, and tensors also work. Video gets sampled at one frame per second by default. Audio should be 16 kHz mono for best results.
EmbeddingGemma 2 scores 14% higher than EmbeddingGemma 1 on MTEB benchmarks. This benchmark specifically measures code understanding capabilities. It adds image, video, and audio retrieval while retaining accuracy. The multilingual text performance matches EmbeddingGemma 1 closely. Engineers can validate these numbers against their own datasets too.
Ready to explore multimodal embeddings? Take a look at the following resources to find out more. Download the Weights directly from Hugging Face. Access them under the Apache 2.0 license for open use. Read more about EmbeddingGemma 2 in the Gemma documentation. Use your favorite development tools like vLLM or Ollama. Run the model efficiently with these popular inference engines.
Adapt and fine-tune for rapid experimentation with Unsloth. Explore efficient fine-tuning options for custom datasets. Try the Instant Media Search on-device demos in the Google AI Edge Gallery. Bring multimodal semantic search to the edge with this new model. Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA. Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs is another related update.
Announcement - Google introduces EmbeddingGemma 2 as a new open-weight multimodal embedding model
Google released a new model called EmbeddingGemma 2. It is designed for search and retrieval tasks. The team at Google created this specific tool. They want developers to use it easily. This model handles many different types of data. It works with text, code, images, video, and audio. You can run it on your own devices. No one owns the code completely. The Apache 2.0 license allows free commercial use. Google made this available for everyone to try.
Architecture - The model uses modular encoders that project into a unified 768-dimensional vector space
The system uses separate parts for different data types. Each part processes its specific input type. All parts feed into one shared backbone layer. This creates a single output space for everything. The final vectors live in a 768-dimensional space. You can load only the pieces you need. Text and code use a base of 270 million parameters. Adding vision support brings the count to 440 million. Including audio pushes it to 570 or 740 million. This flexibility saves memory during startup.
Performance - It outperforms the previous version significantly on code search and technical retrieval tasks
Engineers tested this model against its predecessor. The new version scores much higher on specific tests. Code search accuracy improved by a large margin. Technical retrieval also saw significant gains. The benchmark used is called MTEB. It measures how well models understand code. EmbeddingGemma 2 beats the old version by 14%. Text performance stays very similar to before. This proves the new architecture works better for logic.
Storage Optimization - Matryoshka Representation Learning enables dynamic dimension truncation to save memory
The model supports a technique called Matryoshka Representation Learning. You can shrink the vector size without losing much quality. The original vectors are 768 dimensions long. You can cut them down to 256 or 128 dimensions. This process keeps most of the useful information. Text and code keep about 95% quality at 256 dims. Images, video, and speech keep about 90%. Storage needs drop dramatically with this trick.
Why it matters - This capability reduces costs and complexity for local RAG systems handling mixed media
Local retrieval augmented generation (RAG) systems face big storage bills. They often struggle to index diverse file types. This model solves that problem efficiently. It fits entirely in standard server memory. You do not need massive GPU clusters anymore. Teams can build private search tools easily. The reduced footprint lowers cloud bills significantly. Complex pipelines become much simpler to manage.
How OpenSmartRoute helps
A team routing requests through OpenSmartRoute gains immediate access to EmbeddingGemma 2 without buying new hardware. The router treats this sub-1B model as a single catalogue entry alongside existing options. It scores every candidate on quality, cost, speed and safety before sending traffic.
The system uses Matryoshka Representation Learning to store vectors in just 128 dimensions. This six times storage reduction lowers cloud bills for teams running large vector databases. OpenSmartRoute's input guard spots prompt injection and personal data before any request leaves the network.
Teams can set hard rules so text with personal data stays on an on-premises model. The router ensures a region or cost cap is never crossed during execution. A savings ledger shows what each routed request cost next to the most expensive option.
What to do - Developers can install the model via sentence-transformers and configure modality loading
Start by installing the sentence-transformers library. Use pip with the image and audio flags. Run the command in your terminal window. Then import the SentenceTransformer class. Load the model using the specific ID name. Set config_kwargs to hide unused encoders. This prevents loading heavy parts you do not need. You can run text-only mode for speed. Or load full multimodal support later. Test the similarity function with your data. Validate results against your own datasets.
Mistral AI released a public preview of its largest model, Mistral Large 4. The model features over one trillion parameters and excels in cybersecurity and coding tasks.