Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Abstract
Vision-language models (VLMs) excel at many tasks, yetcontinue to struggle with spatial reasoning, problems where the key in-formation is not directly observable in the input. Many spatial questionsrequire imaginative perception: simulating an unseen viewpoint, tracinga trajectory through an occluded space, or integrating partial views intoa coherent spatial map. Humans naturally support this kind of reason-ing through imagination. Prior work has introduced intermediate visualrepresentations (e.g., visual thoughts, depth, or box tokens), but theseintermediates often refine structure already visible rather than predict-ing the missing spatial structure implied by the evidence. We introduceImaginative Perception Tokens (IPT), intermediate perceptual rep-resentations that externalize what a VLM would perceive under an alter-native spatial configuration while remaining consistent with the observedinput. To study this capability, we formulate three tasks that requireimaginative perception: Perspective Taking (PET), Path Tracing(PT), and Multiview Counting (MVC). For each task, we constructdatasets of →20K examples spanning simulated and real-world settings,paired with ground-truth intermediate imaginations, final answers, andcurated evaluation benchmarks. Using the unified VLM BAGEL [12] asour backbone, IPT supervision improves spatial reasoning across severalsettings and often outperforms textual chain-of-thought training, evenwhen no image is generated at inference time. For example, on MVC,IPT improves accuracy by 3.4% and achieves performance competitivewith strong closed-source models on Path Tracing. We also find thatmixed training with IPT and label-only data can further improve perfor-mance. In contrast, textual chain-of-thought can be detrimental on thesetasks, substantially degrading performance in some cases, highlighting amodality mismatch when forcing spatial computation through language.Overall, IPT provides a principled supervision signal for reasoning overunobserved structure, yielding stronger spatial generalization and a moreinterpretable intermediate aligned with the underlying geometry of thetask. Code will be released at the project page.Path Tracing (11k)Question: As you move from waypoint 1 to 2. This is what you see looking forward from Question: Which object can you seepoint 1: