ModelRefs / Vision Benchmarks — Top AI Models
Vision Benchmarks — Top AI Models
Image understanding, OCR, document parsing, and visual QA. Evaluate visual perception and grounded interpretation.
Overview
Image understanding, OCR, document parsing, and visual QA.
What this category is for: Evaluate visual perception and grounded interpretation.
Benchmarks in this category
- ChartQA — Question answering over real-world charts.
- DocVQA — Document visual question answering on scanned business docs.
- OCRBench — Comprehensive OCR / text-recognition eval for MLLMs.
- AI2D — Diagram understanding across grade-school science textbooks.
- RealWorldQA — Physical-world spatial reasoning from xAI.
- VQAv2 — Open-ended visual question answering test-dev split.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Vision Benchmarks — Top AI Models.