SEERBench: A Spatial Ego-Exo Reasoning Benchmark for MLLMs with a Simple Yet Effective Baseline
Abstract
Humans seamlessly integrate egocentric perceptions with exocentric perspectives to comprehend dynamic spatial positioning. For Embodied AI, mastering this ego-exo spatial reasoning is a fundamental prerequisite for real-world interaction. Although Multimodal Large Language Models (MLLMs) have demonstrated significant progress in general visual reasoning, their capacity for ego-exo spatial reasoning in real-world environments remains largely unexplored. To investigate whether current MLLMs exhibit such human-like spatial cognition, we introduce SEERBench, a rigorous benchmark comprising 1,198 highquality, human-annotated question-answer pairs across 35 diverse realworld Ego-Exo scenarios. SEERBench evaluates models through 10 distinct tasks structured into three progressive cognitive levels: Spatial Perception, Spatial Imagination, and Ego-Exo Collaboration. Extensive evaluations of 19 representative MLLMs reveal a substantial performance gap: even the best-performing frontier models trail human experts by a margin of 46.8%, highlighting severe limitations in current multimodal foundations. To address this, we propose SEER-Map, a training-free, tool-augmented baseline. By explicitly constructing top-down Bird’s-Eye View (BEV) spatial topologies to assist in ego-exo geometric alignment, SEER-Map functions as an Oracle-informed probing baseline lower the threshold of spatial understanding. Codes and data are available here.