CortexVideo: A Semantic-Spatial Dual-Anchor Framework for High-Fidelity fMRI-to-Video Reconstruction
Abstract
Recent advances in neural decoding have enabled the direct reconstruction of visual stimuli from non-invasive brain signals. However, reconstructing continuous video streams remains highly challenging due to the continuous spatial and semantic transformations inherent in dynamic scenes. Most methods rely on a "reconstruct-thendescribe" paradigm, where captions are generated from reconstructed video keyframes, which are highly prone to semantic drift. To overcome these challenges, we propose a novel decoding framework, CortexVideo, to reconstruct video from functional magnetic resonance imaging (fMRI). Inspired by the visual dual-stream hypothesis, CortexVideo introduces dual-guidance decoding strategy. Specifically, we leverage subject-adaptive semantic information to guide video reconstruction, while concurrently integrating keyframe perceptual weights. Experiments demonstrate that CortexVideo provides a cognitive enhancement to existing hierarchical models along two critical dimensions: video semantic understanding and spatial localization. The reconstructed videos show greater fidelity to the visual stimuli in both semantic and spatial aspects, achieving state-ofthe-art results.