Oct 8 edition/Reporting & analysis
ModelsMultimodalInfrastructureBusiness

ModelsArchitectures & capability

Google releases EmbeddingGemma 2, an Apache 2.0 multimodal embedding model small enough for phones

Google's EmbeddingGemma 2 is a 740M-parameter, Apache 2.0 model that embeds text, code, images, video and audio into one vector space and runs on phones. So far, Google has reported its own benchmark scores, and only one small independent test exists.

THE CORE IDEAS4 TAKEAWAYS
01

EmbeddingGemma 2 maps text, code, images, video and audio into a single 768-dimension space. It has an 8K-token context window, and its weights are available under Apache 2.0. The vision and audio encoders are optional modules. Google says memory use ranges from about 191MB for quantized text-only use to about 567MB with every modality loaded on a phone. [3] [5] [6]

02

Simon Willison argues that the open license matters most because stored embeddings tie customers to a model. If a hosted-only embedder is retired, every vector in the index has to be recomputed. Open weights let teams self-host or switch providers and keep their existing indexes. [1]

03

All headline scores come from Google, including an MTEB Code score of 78.68, up from 68.76 for the first version. Leaderboard reviewers flagged problems with Google's results submission. One small hobbyist benchmark found the model matched Qwen3-Embedding-4B on retrieval and ran roughly eight times faster on long inputs. [3] [5] [7] [8]

04

Using the model well takes some extra work. It needs task prefixes on text inputs, and Google recommends bf16 or fp32 precision rather than fp16. In the one independent test, similarity scores fell in a high, narrow band, so thresholds tuned for other models did not carry over, and several long documents overflowed the context window. [5] [8]

WHY IT MATTERS

the open weights allow self-hosting and on-device mixed-media search. Implication: teams could face lower switching risk and send less data to the cloud. However, Google's quality claims have not yet been independently validated.

Executive brief

Google has released EmbeddingGemma 2, a 740M-parameter open-weight model that puts text, code, images, video and audio into one 768-dimension vector space. Google says the full multimodal model runs in about 567MB of RAM on a Pixel 11 Pro (Google). It is licensed under Apache 2.0. Simon Willison argues that the open license matters most because stored embeddings outlive vendors: if a hosted-only model is retired, every stored vector has to be recomputed (Simon Willison). All headline scores come from Google. Google's leaderboard submission is still open, and the only third-party test so far is a small hobbyist benchmark.

What changed and event timeline

  1. EmbeddingGemma 1 paper

    Google described its earlier text-only, lightweight embedder in an arXiv paper. Version 2 is measured against this baseline ().

  2. EmbeddingGemma 2 launches

    Google announced the 740M multimodal model with an 8K context window and Apache 2.0 weights on Hugging Face and Kaggle (;).

  3. Edge tooling ships

    Google AI Edge published LiteRT packaging and latency figures, including 37.3ms per image on a MacBook M5 Pro GPU ().

  4. Leaderboard submission flagged

    Results for three variants were submitted to the MTEB results repository. Reviewers flagged an outdated MTEB version (1.34.7) and a missing run_settings file ().

  5. Willison on lock-in

    On Hacker News, Willison wrote that open weights protect stored vectors even for teams that would rather pay for hosting ().

  6. 6–7 Oct 2026: First hobbyist test

    A homelab retrieval benchmark found the model matching a Qwen3-Embedding-4B control and running much faster on long inputs.

Capabilities and access

  • Model: google/embeddinggemma-2, 740M parameters in total.
  • Input limits: 8,192 tokens, which is about 29 images, 58 video frames at 1 fps, or about 327 seconds of 16kHz mono audio.
  • Output size: 768 dimensions by default. Matryoshka truncation can cut this to 512, 256 or 128.
Read the full section
  • Model: google/embeddinggemma-2, 740M parameters in total. It combines a 270M text backbone with an optional 170M vision encoder and an optional 300M audio encoder (Hugging Face).
  • Input limits: 8,192 tokens, which is about 29 images, 58 video frames at 1 fps, or about 327 seconds of 16kHz mono audio. Google says it covers 100+ languages.
  • Output size: 768 dimensions by default. Matryoshka truncation can cut this to 512, 256 or 128.
  • Runtimes: transformers.js, WebGPU, MLX, vLLM, llama.cpp, Ollama, LM Studio and LiteRT. Google says Model Garden support is coming soon (Google).

Technical analysis for researchers and developers

  • Architecture: Gemma 4-based, with 24 layers, a model dimension of 512, a hidden dimension of 2048, and GQA/MQA attention (Hugging Face).
  • Modular encoders: You can load only the encoders you need. RAM ranges from about 191MB (text only, quantized) to about 567MB (all modalities) (Google Developers Blog).
  • Quantization: INT4 and INT8 quantization-aware training. Truncating from 768 to 128 dimensions alone gives 6x.
Read the full section
  • Architecture: Gemma 4-based, with 24 layers, a model dimension of 512, a hidden dimension of 2048, and GQA/MQA attention (Hugging Face).
  • Modular encoders: You can load only the encoders you need. RAM ranges from about 191MB (text only, quantized) to about 567MB (all modalities) (Google Developers Blog).
  • Quantization: INT4 and INT8 quantization-aware training. Google quotes storage savings as "up to 6x" in one post and "up to 8x" in the other. Truncating from 768 to 128 dimensions alone gives 6x.
  • Implementation details: Text inputs need task prefixes, such as task: search result | query:. Google recommends bf16 or fp32, not fp16.
  • Calibration: In the hobbyist test, scores clustered in a high, narrow band, so similarity thresholds tuned for other models did not transfer (GitHub PR #189).

Claims and evidence

  • MTEB Code 78.68, up from 68.76 for EmbeddingGemma 1 ()
  • MTEB Multilingual 61.36, MIEB 64.64, MMEB Video 50.67, MSEB Audio 69.54 () — Vendor-reported.
  • MAEB score of 49.39 for the audio variant () — Vendor-submitted, not yet accepted.
Read the full section
ClaimStatus
MTEB Code 78.68, up from 68.76 for EmbeddingGemma 1 (Google)Vendor-reported. The leaderboard PR is unmerged (PR #745).
MTEB Multilingual 61.36, MIEB 64.64, MMEB Video 50.67, MSEB Audio 69.54 (Hugging Face)Vendor-reported.
MAEB score of 49.39 for the audio variant (PR #745)Vendor-submitted, not yet accepted.
Best-in-class among sub-1B multimodal embeddersVendor claim with no independent corroboration.
Matched Qwen3-Embedding-4B on top-1/top-5 (19/34 and 28/34 vs 18/34 and 24/34) and was about 8x faster on long inputs (PR #189)Independent, but a single small test (1,280 documents, 34 positives).

Context and prior work

  • EmbeddingGemma 1: Google's earlier text-only embedder, aimed at lightweight on-device use (arXiv). Google says it has about 20M downloads (The Next Web).
  • Gemini Embedding 2: Google's own natively multimodal embedding model, described in a separate arXiv paper (arXiv 2605.27295).
  • Re-embedding costs: Willison recalls that OpenAI once offered to pay for customers re-embedding content when it retired older models.
Read the full section
  • EmbeddingGemma 1: Google's earlier text-only embedder, aimed at lightweight on-device use (arXiv). Google says it has about 20M downloads (The Next Web).
  • Gemini Embedding 2: Google's own natively multimodal embedding model, described in a separate arXiv paper (arXiv 2605.27295).
  • Re-embedding costs: Willison recalls that OpenAI once offered to pay for customers re-embedding content when it retired older models. He argues buyers cannot count on that from every provider (Simon Willison).

Limitations, safety and contested findings

  • Safety tuning: The Next Web reports that no safety tuning was applied (The Next Web).
  • Training data: Data runs to January 2025 and was filtered for CSAM and sensitive data.
  • Benchmark validity: Leaderboard reviewers flagged problems with the MTEB version and the run settings. That leaves Google's leaderboard numbers unvalidated for now (PR #745).
Read the full section
  • Safety tuning: The Next Web reports that no safety tuning was applied (The Next Web).
  • Training data: Data runs to January 2025 and was filtered for CSAM and sensitive data. Google warns that quality varies across languages and drops if task prefixes are omitted (Hugging Face).
  • Benchmark validity: Leaderboard reviewers flagged problems with the MTEB version and the run settings. That leaves Google's leaderboard numbers unvalidated for now (PR #745).
  • Long inputs: Five documents in the homelab test overflowed the context window (PR #189).
  • Storage savings: Google's two posts quote different figures (6x vs 8x).

Business and practitioner implications

  • Lower switching risk: Apache 2.0 weights mean an index built with a hosted version of the model can be moved to self-hosting or another vendor without recomputing the vectors (Simon Willison).
  • One index for every media type: On-device embedding of mixed media allows local search without sending data to the cloud.
  • Before you adopt it: Run your own retrieval evaluation, recalibrate similarity thresholds, enforce task prefixes, and handle inputs longer than 8K tokens explicitly (PR #189).
Read the full section
  • Lower switching risk: Apache 2.0 weights mean an index built with a hosted version of the model can be moved to self-hosting or another vendor without recomputing the vectors (Simon Willison).
  • One index for every media type: On-device embedding of mixed media allows local search without sending data to the cloud. The Next Web ties this to European data-protection concerns (The Next Web).
  • Before you adopt it: Run your own retrieval evaluation, recalibrate similarity thresholds, enforce task prefixes, and handle inputs longer than 8K tokens explicitly (PR #189).
FOLLOW THE EVIDENCE

The source trail.

Sources (10)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief