VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context
Abstract
Large Vision Language Models (LVLMs) have achieved re-markable success on vision–language tasks, yet fine-grained perceptionover high-resolution images and long-context videos remains challenging.As the number of visual tokens increases, the visual attention sink phe-nomenon becomes increasingly severe, causing irrelevant tokens to absorba disproportionate amount of attention mass. Recent approaches attemptto mitigate this issue by explicitly predicting bounding boxes or temporalspans and re-encoding the cropped visual regions. Such methods dependon unreliable numeric localization in the discrete token space and in-cur significant computational overhead due to additional forward passes.In this work, we propose VisReflect, a simple yet effective frameworkthat improves fine-grained perception in long visual contexts through la-tent visual reflection. Instead of decoding intermediate predictions intodiscrete tokens, the model generates continuous visual reflection thatrepresents question-relevant visual features in the latent space. Thesereflections selectively emphasize salient regions or frames, guiding at-tention towards relevant visual tokens within a single forward pass. Weconduct comprehensive evaluations on challenging high-resolution im-age benchmarks, including BLINK, V∗ , and HRBench-4K/8K, as wellas video understanding benchmarks such as MVBench, VideoMME, andMLVU. Our method consistently improves over strong baselines, achiev-ing gains of 4.1% on image benchmarks and 1.8% on video benchmarks.Compared with zooming-based methods, our model achieves compara-ble performance while reducing inference time by roughly 44% on videounderstanding.