ModelRefs / Retrieval Benchmarks — Top AI Models
Retrieval Benchmarks — Top AI Models
Embedding and retrieval benchmarks for semantic search and multilingual RAG. Evaluate retrieval-oriented representation quality before application-specific RAG testing.
Overview
Embedding and retrieval benchmarks for semantic search and multilingual RAG.
What this category is for: Evaluate retrieval-oriented representation quality before application-specific RAG testing.
Benchmarks in this category
- MTEB — Massive Text Embedding Benchmark covering retrieval and other embedding tasks across datasets and languages.
- MIRACL — Multilingual information retrieval benchmark spanning 18 languages and diverse topic sets.
- MKQA — Cross-lingual question retrieval benchmark built from Multilingual Knowledge Questions & Answers across 26 languages.
- MLDR — Multilingual long-document retrieval benchmark introduced with M3-Embedding for retrieval across 13 languages.
- BrowseComp Long Context — OpenAI long-context question-answering benchmark with relevant search results embedded in inputs up to hundreds of thousands of tokens.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Retrieval Benchmarks — Top AI Models.