UniGeo: Unifying Geometric Constraints for Camera-Controllable Image Editing via Video Priors
Abstract
Camera-controllable view synthesis aims to synthesize novelviews of a given scene under varying camera poses while strictly preservingcross-view geometric consistency. However, existing methods typicallyrely on fragmented geometric guidance, such as only injecting point cloudsat the representation level despite models containing multiple levels, andare mainly based on image diffusion models that operate on discreteview mappings. These two limitations jointly lead to geometric drift andstructural degradation under continuous camera motion. We observethat while leveraging video models provides continuous viewpoint priorsfor camera-controllable view synthesis, they still struggle to form stablegeometric understanding if geometric guidance remains fragmented. Tosystematically address this, we inject unified geometric guidance across thethree levels that jointly determine the generative output: representation,architecture, and loss function. To this end, we propose UniGeo, a novelcamera-controllable editing framework. Specifically, at the representationlevel, UniGeo incorporates a frame-decoupled geometric reference injectionmechanism to provide robust cross-view geometry context. Furthermore,at the architecture level, it introduces a geometric anchor attention toalign multi-view features, and at the loss function level, it proposes atrajectory-endpoint geometric supervision strategy to explicitly reinforcethe structural fidelity of target views. Experiments across multiple publicbenchmarks, encompassing both extensive and limited camera motionsettings, demonstrate that UniGeo significantly outperforms existingmethods in visual quality and geometric consistency.