ModelRefs / Multimodal Benchmarks — Top AI Models
Multimodal Benchmarks — Top AI Models
Cross-modal reasoning across text, image, audio, and video. Benchmark models that operate across modalities at once.
Overview
Cross-modal reasoning across text, image, audio, and video.
What this category is for: Benchmark models that operate across modalities at once.
Benchmarks in this category
- MMMU — Massive Multi-discipline Multimodal Understanding across 30 subjects.
- MMMU-Pro — Harder, contamination-resistant successor to MMMU.
- MathVista — Math reasoning over visual contexts (charts, figures, geometry).
- Video-MME — Comprehensive video understanding eval across 6 domains.
- EgoSchema — Long-form first-person video question answering.
- MLVU — Multi-task long video understanding (3 min – 2 hr clips).
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Multimodal Benchmarks — Top AI Models.