ModelRefs / Multimodal Benchmarks — Top AI Models

Multimodal Benchmarks — Top AI Models

Cross-modal reasoning across text, image, audio, and video. Benchmark models that operate across modalities at once.

Overview

Cross-modal reasoning across text, image, audio, and video.

What this category is for: Benchmark models that operate across modalities at once.

Benchmarks in this category

  • MMMU — Massive Multi-discipline Multimodal Understanding across 30 subjects.
  • MMMU-Pro — Harder, contamination-resistant successor to MMMU.
  • MathVista — Math reasoning over visual contexts (charts, figures, geometry).
  • Video-MME — Comprehensive video understanding eval across 6 domains.
  • EgoSchema — Long-form first-person video question answering.
  • MLVU — Multi-task long video understanding (3 min – 2 hr clips).

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Multimodal Benchmarks — Top AI Models.