← Blog

28 March 2026

Multilingual multimodal RAG: why your AI ignores your non-English documents

Multilingual RAG, comparison of embedding models

Most RAG systems running today have two blind spots: they understand only English and they see only text.

What the usual benchmarks do not test

The MTEB leaderboard tests one thing: searching English text in a base of English text. In practice, companies work with documents in Dutch, French or German, with images and with tables.

Milvus built the CCKM benchmark, which tests multilingual search, multimodal search and search inside long documents (32,000 characters).

Multilingual RAG results

  • Gemini Embedding 2: 99.7%
  • Qwen3-VL-2B (open source): 98.8%
  • OpenAI text-embedding-3-large: 96.7%
  • Lightweight models (nomic-embed-text, mxbai-embed-large): 12-15%
  • On idioms: 3%

Only Gemini scored perfectly on aligning Chinese and English idioms.

Multimodal search results

  • Qwen3-VL-2B: 94.5%
  • Gemini: 92.8%

The modality gap is what decides it: Qwen (0.25) against Gemini (0.73).

What this means outside the English-speaking world

The choice of embedding model decides whether your AI understands your documents or ignores them. AS3P uses Milvus for BrainDup, an open source vector database that handles multilingual and multimodal content natively.

There is more than English in the world of AI. And there is more than text.

Source: Milvus, CCKM benchmark.