AI product Open source
PerceptionBench is a benchmark for evaluating atomic visual perception in multimodal large language models (MLLMs). It isolates perception from reasoning and domain knowledge by defining ten atomic perceptual capabilities, including counting, depth, localization, OCR, and perception-related hallucination, based on an error taxonomy derived from failures across existing benchmarks. The released dataset contains 3,000 verified open-ended questions with short, uniquely determined answers; 1,800 are decomposed from failures in source benchmarks and 1,200 are newly authored using supplementary images. The repository provides the benchmark's code and links to its paper and dataset.
1 use taken from transcripts — each links to the moment in the video.
Evaluates whether multimodal models can perceive visual information before reasoning about it. It contains 3,000 verified open-ended questions across ten skills, including counting, depth, localization, OCR, and resistance to hallucination.
1 in the library.