AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
Abstract
Diffusion models have significantly advanced video super-resolution (VSR) but remain largely constrained to fixed upsamplingscales. Conversely, while coordinate-based arbitrary-scale VSR methodsoffer scale flexibility, they inherently suffer from severe over-smoothingat large scaling factors. Integrating generative priors with continuous de-coding is promising but currently hindered by severe temporal flickeringcaused by the stochasticity of diffusion sampling. To address this, wepropose AVSR-Diff (Arbitrary-scale Video Super-Resolution with Dif-fusion), a novel decoupled framework that separates scale-agnostic la-tent denoising from continuous coordinate rendering, effectively avoid-ing computationally heavy resolution-specific sampling. Our approachintroduces a Temporally-Gated Feature Recurrence (TGFR) module toextract strictly aligned, temporally consistent latent priors. Furthermore,we design a continuous video VAE decoder incorporating a Scale-AwareFourier Refinement (SAFR) module to dynamically adapt frequencycomponents to any target scale. Extensive experiments demonstrate thatAVSR-Diff consistently preserves high-frequency details and strong tem-poral stability across various scales, surpassing state-of-the-art arbitrary-scale baselines. Remarkably, our framework outperforms recent fixed-scale generative models even on their native resolution.