October 9, 2026

Best Embedding Models in 2026: EmbeddingGemma 2, Perplexity, Qwen3 and More Compared

1

A practical comparison of the best embedding models in 2026, including Google’s new EmbeddingGemma 2 and Perplexity’s pplx-embed-v2-late, with tips on choosing one for RAG and semantic search.

Abstract network of connected dots, illustrating how the best embedding models map text into vector space

Photo by <a href="https://unsplash.com/@choys_?utm_source=WP+Agent&utm_medium=referral">Conny Schneider</a> on <a href="https://unsplash.com/?utm_source=WP+Agent&utm_medium=referral">Unsplash</a>

Choosing among the best embedding models got harder this week. On 6 October 2026 Google released EmbeddingGemma 2, a small open model that embeds text, images, audio and video in one vector space, and a day later Perplexity published its pplx-embed-v2-late models for multimodal document search. This guide explains what embedding models do, compares the strongest open and hosted options for retrieval-augmented generation (RAG) and semantic search, and gives you a simple way to choose and test one on your own data.

Abstract network of connected dots, illustrating how the best embedding models map text into vector space
Photo by Conny Schneider on Unsplash

Specifications below were checked on 8 October 2026 against Google’s EmbeddingGemma 2 announcement, Perplexity’s research post, model cards on Hugging Face and the OpenAI and Mistral documentation. Benchmark scores are the developers’ own figures unless stated otherwise; they are useful for shortlisting, but the only result that really matters is how a model performs on your documents.

What Is an Embedding Model?

An embedding model turns a piece of content, such as a sentence, a PDF page or an image, into a list of numbers called a vector. Content with similar meaning ends up with similar vectors, so a computer can find related items by measuring the distance between them. That makes embeddings the engine behind:

  • RAG: finding the right passages to feed a chatbot so it answers from your own documents.
  • Semantic search: matching “refund policy” to a page titled “returns and exchanges”.
  • Recommendations and clustering: grouping similar tickets, products or articles.
  • Duplicate detection and classification.

Document chat tools like the one in our NotebookLM guide rely on this kind of retrieval behind the scenes. When you build your own system, choosing the embedding model is one of the most important decisions you make.

Best Embedding Models in 2026 at a Glance

Model Type Size Modalities Licence / access Best for
EmbeddingGemma 2 (Google) Open weights 270M–740M Text, code, images, audio, video Apache 2.0 On-device and local multimodal search
pplx-embed-v2-late (Perplexity) Open weights 0.6B and 9B Text, images, PDF pages Hugging Face; check the model card for licence Visual document and PDF retrieval
Qwen3-Embedding (Alibaba) Open weights 0.6B, 4B, 8B Text and code Apache 2.0 Multilingual text RAG
OpenAI text-embedding-3 Hosted API Small and large Text Paid API Quick, managed text embeddings
Mistral Embed / Codestral Embed Hosted API Not published Text; code Paid API Teams already on Mistral’s EU platform

1. EmbeddingGemma 2: Best Small Multimodal Model

Google DeepMind released EmbeddingGemma 2 on 6 October 2026. It builds on Gemma 4 and, according to Google, maps combinations of text, images, audio and video into a single shared embedding space.

  • Size: 740 million parameters in total. Text-only use needs as little as 270M, with an optional 170M vision encoder and 300M audio encoder.
  • Context: 8,192 tokens, four times the first version. Google says that fits about 29 images, 58 video frames or 5.5 minutes of audio.
  • Dimensions: 768 by default, which can be shortened to 512, 256 or 128 with Matryoshka Representation Learning, cutting storage by up to six times.
  • Licence: Apache 2.0, so commercial use is allowed.
  • Where to get it: weights on Hugging Face and Kaggle, plus support in Ollama, llama.cpp, MLX, vLLM, LiteRT and other tools. A hosted option in Google’s enterprise Model Garden is “coming soon”.

MarkTechPost reports an MTEB multilingual score of 61.36 and an MTEB Code score of 78.68 (up from 68.76 for version 1), with quantised text-only weights using roughly 191 MB of RAM on a phone. Choose it if you want a small, permissively licensed model that runs locally and handles more than text.

Multicoloured straws arranged in a circle, representing multimodal embedding models that combine text, images and audio
Photo by Mitul Grover on Unsplash

2. Perplexity pplx-embed-v2-late: Best for PDFs and Visual Documents

Perplexity, better known for its AI answer engine (see our Perplexity vs ChatGPT comparison), has released the new pplx-embed-v2-late family, which takes a different approach called late interaction (the ColBERT style). Instead of squeezing a whole document into one vector, it keeps a small 128-dimensional vector for every token and compares queries token by token. That tends to improve accuracy on complex documents, at the cost of a larger index.

  • Sizes: a 0.6B model aimed at laptops and edge devices, and a 9B model for maximum accuracy. Both share one embedding space, so you can index with the 9B and query with the 0.6B.
  • Modalities: text, images and rendered PDF pages. Pages are encoded as images, so no OCR step is needed.
  • Training: distilled from an 18B teacher on 186 million query-document pairs in 46 languages, according to Perplexity.
  • Availability: weights on Hugging Face; Perplexity says API support will follow later.

Perplexity reports 92.4% on the MADQA document question-answering benchmark for the 9B model and 90.1% for the 0.6B, and says the 9B leads its 72-task domain-specific text retrieval average at 81.3%. It also notes that it does not lead the ViDoRe v3 image retrieval benchmark, where another model scores higher. Perplexity’s page does not state a licence; MarkTechPost reports MIT, but confirm on the Hugging Face model card before commercial use. Choose it if your knowledge base is full of PDFs, slides or scanned pages.

3. Qwen3-Embedding: Best Open Model for Multilingual Text

Alibaba’s Qwen3-Embedding series remains a strong text-only choice. According to its Hugging Face model card it comes in 0.6B, 4B and 8B sizes, supports more than 100 languages including programming languages, uses the Apache 2.0 licence and offers a 32K-token context window on the 0.6B model. The 0.6B model outputs up to 1,024 dimensions (adjustable down to 32), while the 4B and 8B use 2,560 and 4,096.

The model card says the 8B model ranked first on the MTEB multilingual leaderboard with a score of 70.58 as of June 2025. That ranking is now more than a year old, so check the current leaderboard, but Qwen3-Embedding is still a sensible default for multilingual text RAG in German, French, Spanish, Polish and other European languages. There are matching reranker models too.

4. OpenAI text-embedding-3: Easiest Hosted Option

If you want embeddings without running anything yourself, OpenAI’s API offers text-embedding-3-small (1,536 dimensions by default) and text-embedding-3-large (3,072 dimensions), both with an 8,192-token input limit and a dimensions parameter to shorten vectors. OpenAI’s guide prices them in “pages per dollar”: 62,500 pages for small and 9,615 for large, assuming about 800 tokens per page. That works out to roughly $0.02 and $0.13 per million tokens respectively; check OpenAI’s pricing page for current rates.

5. Mistral Embed and Codestral Embed: For EU-Focused Teams

Mistral’s model overview lists two embedding models: Mistral Embed for general text and Codestral Embed for code. They are hosted on Mistral’s platform alongside its chat models, which is convenient if you already use Mistral for data-residency reasons. Mistral does not publish sizes or open weights for these models. If you are considering Mistral’s newest chat model too, see our Mistral Large 4 guide.

Other models worth knowing

Google also offers a hosted Gemini Embedding 2 model, which Perplexity used as a comparison point in its own benchmarks. Community favourites such as BGE-M3 and nomic-embed-text appear regularly in comparisons and are worth including in your tests.

How to Choose the Best Embedding Model for RAG

Library bookshelves, representing document retrieval with embedding models for RAG
Photo by Emil Widlund on Unsplash
  1. Start with your data type. Plain text points to Qwen3-Embedding, EmbeddingGemma 2 or a hosted API. Scanned PDFs, slides and screenshots point to pplx-embed-v2-late. Mixed media (images, audio, video) points to EmbeddingGemma 2.
  2. Decide where it must run. If data cannot leave your servers or devices, which is common under GDPR, choose an open-weight model. Our guide on how to run an LLM locally covers the tools, such as Ollama and llama.cpp, that also serve embedding models.
  3. Check language coverage. For European languages, prioritise multilingual models and test with real queries in each language.
  4. Balance dimensions and storage. Bigger vectors can be more accurate but cost more to store and search. Models with Matryoshka support let you shrink vectors with only a small loss in quality.
  5. Check the licence. Apache 2.0 and MIT are commercially friendly; always confirm on the official model card.
  6. Benchmark on your own documents. Build 30–50 real questions with known correct passages, and measure how often each model puts the right passage in its top five results.

How to Try an Embedding Model Locally

The quickest way to experiment is with Ollama, which can serve embedding models on your computer. EmbeddingGemma 2 is listed in Ollama’s library; check the exact tag on the model page, then pull it and request an embedding from the local API:

ollama pull embeddinggemma-2

curl http://localhost:11434/api/embed -d '{
  "model": "embeddinggemma-2",
  "input": "What is our refund policy for damaged items?"
}'

Developers working in Python can use the sentence-transformers library, which Google lists as supported (version 6.1 or later):

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")
vectors = model.encode([
    "How do I reset my password?",
    "Steps to change your account password",
])
print(vectors.shape)

Read each model card before going to production: some models expect special prompt prefixes for queries and documents, and using them can noticeably improve results. Store the vectors in a vector database (Google lists Qdrant among supported options) and you have the retrieval half of a RAG system.

Open vs Hosted Embedding Models

Open-weight models Hosted APIs
Data location Stays on your hardware Sent to the provider
Cost Your hardware and electricity Pay per token
Setup More work One API call
Model changes You control versions Provider may retire models
Examples EmbeddingGemma 2, pplx-embed-v2-late, Qwen3-Embedding OpenAI text-embedding-3, Mistral Embed

One practical warning: vectors from different models are not compatible. If you switch embedding models later, you must re-embed your whole collection. That is a good reason to test carefully before indexing millions of documents. If you already use a cheap hosted chat model, our DeepSeek API guide shows how retrieval can be paired with low-cost generation.

Frequently Asked Questions

What is the best embedding model in 2026?

There is no single winner. For small, local and multimodal use, EmbeddingGemma 2 stands out. For PDFs and visual documents, Perplexity’s pplx-embed-v2-late reports the strongest results. For multilingual text, Qwen3-Embedding is a proven open option, and OpenAI’s text-embedding-3 models are the simplest hosted choice.

What is the best embedding model for RAG?

It depends on your documents. Text-heavy knowledge bases suit Qwen3-Embedding or EmbeddingGemma 2; PDF-heavy ones suit pplx-embed-v2-late. Always test on your own questions.

Which embedding models are free and open source?

EmbeddingGemma 2 and Qwen3-Embedding use Apache 2.0. Perplexity’s pplx-embed-v2-late weights are on Hugging Face; check the model card for licence terms.

What is the best embedding model for Ollama?

EmbeddingGemma 2 is a strong fit because of its small size and official Ollama support. Check the Ollama library for other embedding models and their tags.

Can I mix embeddings from different models?

No. Each model creates its own vector space. Queries and documents must be embedded with the same model, or with models designed to share a space, like the two pplx-embed-v2-late sizes.

What is the best multilingual embedding model?

Qwen3-Embedding and EmbeddingGemma 2 both support more than 100 languages. Test with real queries in each language your users speak.

Conclusion

The best embedding models in 2026 are more capable and more varied than ever. EmbeddingGemma 2 brings multimodal search to phones and laptops under an Apache 2.0 licence, Perplexity’s pplx-embed-v2-late raises the bar for PDF and document retrieval, Qwen3-Embedding remains a strong multilingual workhorse, and hosted APIs from OpenAI and Mistral offer the fastest route to production. Shortlist two or three, test them on your own documents and questions, and choose the one that finds the right answer most often for a cost and setup you can live with.

Sources: Google – EmbeddingGemma 2; MarkTechPost on EmbeddingGemma 2; Perplexity – Multimodal embeddings beyond a single vector; MarkTechPost on pplx-embed-v2-late; Qwen3-Embedding model card; OpenAI embeddings guide; Mistral models overview.

1 thought on “Best Embedding Models in 2026: EmbeddingGemma 2, Perplexity, Qwen3 and More Compared”

Leave a Reply

Your email address will not be published. Required fields are marked *