Vision LLM
Benchmark

Extracting information from images.

VLMBench v1: an early public benchmark for VLM document reasoning.

7 models
550 questions
115,500 graded cells
v1 · 2026
VLMBench per-model results: Claude Opus 5 at 99.99 percent acceptable, with per-question run grids below
Explore the full results →

Each model

  • same image
  • same prompt
  • same question
×30

Analyze

  • Track consistency and accuracy
  • Observe the answer / solution space
  • Map error types

Results

Acceptable answers (exact + materially correct) out of 16,500 graded cells per model.

Recommend a model

Error Types

Errors were classified, here is the breakdown:

Accuracy and consistency

Accuracy is far and away most important.

Suggest