Most RAG systems running today have two blind spots: they understand only English and they see only text.
What the usual benchmarks do not test
The MTEB leaderboard tests one thing: searching English text in a base of English text. In practice, companies work with documents in Dutch, French or German, with images and with tables.
Milvus built the CCKM benchmark, which tests multilingual search, multimodal search and search inside long documents (32,000 characters).
Multilingual RAG results
- Gemini Embedding 2: 99.7%
- Qwen3-VL-2B (open source): 98.8%
- OpenAI text-embedding-3-large: 96.7%
- Lightweight models (nomic-embed-text, mxbai-embed-large): 12-15%
- On idioms: 3%
Only Gemini scored perfectly on aligning Chinese and English idioms.
Multimodal search results
- Qwen3-VL-2B: 94.5%
- Gemini: 92.8%
The modality gap is what decides it: Qwen (0.25) against Gemini (0.73).
What this means outside the English-speaking world
The choice of embedding model decides whether your AI understands your documents or ignores them. AS3P uses Milvus for BrainDup, an open source vector database that handles multilingual and multimodal content natively.
There is more than English in the world of AI. And there is more than text.
Source: Milvus, CCKM benchmark.