CabinSI: Omni-Cabin Spatial Reasoning through Explicit Visual Cognitive Maps
Abstract
Intelligent cabin is gradually evolving from passive voicebased interfaces into multimodal collaborative systems supported by multi-camera and multi-sensor setups, where spatial reasoning becomes a core capability for enabling proactive perceptual interaction. Unlike traditional single-scene spatial tasks, cabin environments involve both interior and exterior vehicle regions, multiple occupants and traffic participants, and inherently require cross-view observations. To fill this research gap, we introduce CabinSI, a benchmark for evaluating spatial intelligence in cross-cabin environments, consisting of two components: RelCabin for relational reasoning tasks and RefCabin for spatial referring localization tasks. The benchmark is built upon real-world captured multi-view cabin data and systematically covers in-cabin, out-of-cabin, and cross-cabin scenarios under single-view, multi-view, and cross-cabinview settings. Furthermore, we propose a cognitive-map-based framework that projects multi-view observations onto a normalized top-down plane and constructs an explicit spatial graph as model input, allowing multimodal large language models to focus on structured reasoning. Experimental results demonstrate that such explicit spatial representation significantly improves the stability of cross-view reasoning in MLLMs. Data and code is available at CabinSI.