IC-World: In-Context Generation for Shared World Modeling
Abstract
Video-based world models have recently garnered increas-ing attention for their ability to synthesize diverse and dynamic vi-sual environments. In this paper, we focus on shared world modeling,where a model generates multiple videos from a set of input images,each representing the same underlying world in di!erent camera poses.We propose IC-World, a novel generation framework, enabling paral-lel generation for all shared world input images via activating the in-herent in-context generation capability of large video models. We fur-ther finetune IC-World via reinforcement learning, Group Relative Pol-icy Optimization, together with two proposed novel reward models toenforce scene-level geometry consistency and object-level motion consis-tency among the set of generated videos. Extensive experiments demon-strate that IC-World substantially outperforms state-of-the-art methodsin both geometry and motion consistency. To the best of our knowledge,this is the first work to systematically explore the shared world modelingproblem with video-based world models.