DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues
Abstract
Recent advancements in Multimodal Large Language Models(MLLMs) have demonstrated impressive fine-grained perception capabil-ities. However, existing benchmarks predominantly rely on explicit tex-tual cues or low-resolution inputs, failing to evaluate a model’s ability toautonomously perceive implicit visual cues in high-resolution. To bridgethis gap, we introduce DiCoBench, a comprehensive, multi-image high-resolution benchmark designed for cross-image fine-grained perception.DiCoBench consists of 765 meticulously curated samples categorizedinto two progressive tracks: Differential Visual Cues and Commonal-ity Visual Cues, covering 8 distinct perception tasks. By formulatingthe benchmark as a multiple-choice question task and utilizing high-resolution imagery (approaching 2K), we eliminate evaluation metricbias and pose a substantial challenge to current state-of-the-art MLLMs.Our extensive evaluation of 18 diverse MLLMs reveals a striking perfor-mance gap compared to human accuracy (98.3%), with top-performingmodels struggling significantly with micro-scale detail capture. We be-lieve DiCoBench will serve as a challenging testbed to drive future re-search in autonomous, high-resolution multi-image perception. Dataset:https://huggingface.co/datasets/oking0197/DiCoBench