GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
Abstract
Visual generation models can produce photorealistic videosyet violate multiview geometry, exhibiting non-rigid deformations andocclusion inconsistencies (e.g., hallucinated content in disoccluded re-gions). These failures hinder downstream applications such as video worldmodel development and 3D asset creation, and are poorly captured byexisting metrics. We introduce GeCo, a geometry-grounded metric forjointly detecting geometric deformation and occlusion-inconsistency ar-tifacts in static scenes. By fusing residual motion and depth priors, GeCoproduces interpretable, dense consistency maps that localize these arti-facts. Using GeCo, we systematically benchmark recent video generationmodels, revealing common geometric failure modes. We further showthat GeCo provides an actionable signal by applying it as a training-free guidance loss that substantially reduces geometric artifacts duringgeneration.