ModelRefs / Vision Benchmarks — Top AI Models

Vision Benchmarks — Top AI Models

Image understanding, OCR, document parsing, and visual QA. Evaluate visual perception and grounded interpretation.

Overview

Image understanding, OCR, document parsing, and visual QA.

What this category is for: Evaluate visual perception and grounded interpretation.

Benchmarks in this category

  • ChartQA — Question answering over real-world charts.
  • DocVQA — Document visual question answering on scanned business docs.
  • OCRBench — Comprehensive OCR / text-recognition eval for MLLMs.
  • AI2D — Diagram understanding across grade-school science textbooks.
  • RealWorldQA — Physical-world spatial reasoning from xAI.
  • VQAv2 — Open-ended visual question answering test-dev split.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Vision Benchmarks — Top AI Models.