Formerly Janakpur Engineering College (JEC)Affiliated to Tribhuvan University

Google releases EmbeddingGemma 2, a small open model for search

EmbeddingGemma 2 maps text, code, images, video and audio into one vector space, runs on phones, and is free under the Apache 2.0 licence.

BCTBEI

On 6 October 2026, Google DeepMind released EmbeddingGemma 2, a free open model that turns text, code, images, video and audio into lists of numbers in one shared space, the Google Developers Blog announced. The full model has 740 million parameters, and the text-only part can run in about 191 megabytes of memory on a phone.

  • 768dimensions in the shared vector space for all five content types
  • 740 millionparameters in the full multimodal model
  • 8,192tokens of context, four times the first version
  • 20 milliondownloads of the first EmbeddingGemma, Google says
  • 6xless storage when vectors are cut from 768 to 128 numbers

What happened

EmbeddingGemma 2 is an embedding model. It does not write text or answer questions. Instead, it turns a piece of content into a vector, a list of numbers that represents its meaning. Google released it under the Apache 2.0 licence, which allows free use, including commercial use. The weights are available on Hugging Face and Kaggle, according to Unite.AI, and Google's guide shows how to run it in a few lines of Python.

The model maps five kinds of content, text, code, images, video and audio, into a single space of 768 dimensions. This means a text search can find a matching photo or sound recording, because all of them become comparable vectors. Unite.AI reports that the model is built on Google's Gemma 4 architecture and was announced by Google DeepMind engineers Sahil Dua and Henrique Schechter Vera.

The first EmbeddingGemma came out in 2025 and handled text only. Google says it has passed 20 million downloads. According to Google, the new version scores 14 percent higher on a standard code search test, called MTEB Code, while keeping the same quality on text in many languages. The context window, the amount of input it can read at once, is 8,192 tokens, four times larger than the first version.

The engineering behind it

The model is modular. The text and code part has 270 million parameters. A vision part adds 170 million, and an audio part adds 300 million. A developer loads only the parts needed: 270 million for text, 440 million for text and images, 570 million for text and audio, or 740 million for everything. All four setups come from the same file and produce vectors in the same space, so they can be mixed in one search index.

Search with embeddings works in a general way. Each document is turned into a vector and stored. A user's question is also turned into a vector. The system then finds stored vectors that point in a similar direction, usually by measuring the angle between them. Content with similar meaning ends up close together. Retrieval-augmented generation, or RAG, uses this step to find relevant documents before a chatbot writes its answer.

EmbeddingGemma 2 uses a method called Matryoshka Representation Learning, named after Russian nesting dolls. The most important information is packed into the first numbers of each vector, so a developer can cut it from 768 to 512, 256 or 128 numbers. Google says storing a million full vectors takes about 1.5 gigabytes, and about 250 megabytes at 128 numbers. The cost is quality: at 128 numbers, image, video and speech search drops to around 75 percent.

All inputs share the 8,192-token budget at fixed rates. Unite.AI reports 280 tokens per image, 140 tokens per video frame and 25 tokens per second of audio. That allows up to 29 images, 58 video frames or about 5.5 minutes of audio in one input. Google's guide says video is sampled at one frame per second by default, and audio should be supplied at 16 kilohertz in a single channel.

Limits the makers state

Unite.AI, citing the model card, reports that the model supports more than 100 languages, but its performance may not be equal across them. The training text covered more than 140 languages, with data up to January 2025. The model has not been through safety tuning after training, so Google leaves filtering and fairness checks to the developers who build on it.

The model card also warns about number formats. The model's internal values can be too large for the 16-bit float format, which can produce broken or silently wrong vectors, so Google recommends bfloat16 or 32-bit floats. Leaving out the recommended short task instruction before text also lowers the quality of the vectors. These are the kinds of details that decide whether a student project works or fails quietly.

What it means in Nepal

The sources do not mention Nepal or Nepali, so any use for Devanagari text has to be tested first. Google does not publish results for each language. A student who wants to search Nepali documents should build a small test set of questions and correct answers, and measure how often the model finds the right document. That test is itself a useful piece of engineering work. It also shows whether a shorter vector, such as 256 numbers, is good enough for the task, which saves storage on a small server.

The practical point is cost. A model that runs on an ordinary laptop, with no paid service and a licence that allows free use, makes search and RAG projects possible for students without special hardware. Google says it works with common tools such as Hugging Face Transformers, sentence-transformers, Ollama and llama.cpp. Examples of student projects include a search tool over lecture notes, past exam questions or a library catalogue, with the vectors stored in a database.

What to study if this interests you

Database Management System, ENCT 301, in the fifth semester of BCT, teaches how data is stored, indexed and queried, which is the base for vector search, and the course has a full guide on this site. Artificial Intelligence, ENCT 351 in the sixth semester of BCT, also with a full guide on this site, and ENCT 305 in the fifth semester of BEI, cover the neural networks that produce embeddings. Foundation of Data Science, ENCT 202, in the third semester of BCT, introduces vectors, similarity and evaluation.

Words in this story

Embedding
A list of numbers that represents the meaning of a piece of content, so similar content gets similar numbers.
Retrieval-augmented generation
A method where a system first searches for relevant documents and then gives them to a language model to write an answer.
Vector database
A database built to store embeddings and quickly find the ones closest to a query.
Apache 2.0 licence
An open licence that lets anyone use, change and share software, including for commercial products.

Where this comes from

Written in our own words; no sentence is copied from these reports. Researched with AI assistance on 11 October 2026; no member of faculty has reviewed it yet. If you spot a mistake, call 01-5091616 and we will correct it and say so.

Next story: 1 trillionMistral previews Large 4, a trillion-parameter open-weight model

Last reviewed by Imperial College of Engineering. Written 11 October 2026 from the sources above.