BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
Abstract
Recent advances in generative video models have signifi-cantly improved visual realism in video object removal, yet evaluationprotocols still focus on masked-region fidelity, treating removal as lo-cal inpainting. In real scenes, object removal is a causal intervention:eliminating an object also requires removing its induced physical ef-fects, such as shadows, reflections, illumination changes, translucency,and dynamic traces. Existing benchmarks lack aligned clean referencesor remain limited to simplified synthetic settings, preventing systematicevaluation of causal consistency. We introduce BeyondMasks, a pairedbenchmark for causally consistent video object removal, consisting oftemporally aligned synthetic and real-world video pairs with clean back-ground references. The dataset spans diverse photometric, geometric,volumetric, and dynamic interactions, and supports both mask-basedand instruction-driven editing. We further propose CORE, a structuredvision–language model-based evaluation protocol that jointly measuresobject disappearance and after-effect consistency, aligning more closelywith human judgments than existing metrics. Benchmarking state-of-the-art methods reveals systematic failures in removing secondary physicaleffects despite high masked-region fidelity, exposing a gap between visualplausibility and causal correctness. BeyondMasks reframes video objectremoval as causal scene consistency rather than local reconstruction andprovides a unified framework for its evaluation.