Structured-Noise Masked Modeling for Video, Audio and Beyond
Abstract
Masked modeling has emerged as a robust self-supervisedlearning framework. However, most methods rely on random masking,which disregards the structural properties of different data modalities. Toalign with the spatiotemporal and spectral characteristics of video andaudio data, we introduce a structured noise-based masking approach.By filtering white noise into different color noise distributions, we gen-erate structured masks that capture modality-specific patterns withoutrequiring handcrafted heuristics or access to the data. Our approach en-hances masked video and audio modeling frameworks without any addi-tional computational cost. Experiments show that structured noise mask-ing consistently outperforms random masking, underscoring the value ofmodality-aware masking strategies for representation learning.