AnyView: Synthesizing Any Novel View in Dynamic Scenes
Abstract
Modern generative video models excel at producing convinc-ing, high-quality outputs, but struggle to maintain multi-view and spa-tiotemporal consistency in highly dynamic real-world environments. Inthis work, we introduce AnyView, a diffusion-based video generationframework for dynamic view synthesis with minimal inductive biasesor geometric assumptions. We leverage multiple data sources with var-ious levels of supervision, including monocular (2D), multi-view static(3D) and multi-view dynamic (4D) datasets, to train a generalist spa-tiotemporal implicit representation capable of producing zero-shot novelvideos from arbitrary camera locations and trajectories. We evaluateAnyView on standard benchmarks, showing competitive results with thecurrent state of the art, and propose AnyViewBench, a challengingnew benchmark tailored towards extreme dynamic view synthesis in di-verse real-world scenarios. In this more dramatic setting, we find thatmost baselines drastically degrade in performance, as they require signif-icant overlap between viewpoints, while AnyView maintains the abilityto produce realistic, plausible, and spatiotemporally consistent videoswhen prompted from any viewpoint.