Grounding World Simulation Models in a Real-World Metropolis
Abstract
What if a world simulation model could render not an imag-ined environment but a city that actually exists? Prior generative worldmodels synthesize visually plausible yet artificial environments by imag-ining all content. We present Seoul World Model (SWM), a city-scale world model grounded in the real city of Seoul. SWM anchors au-toregressive video generation through retrieval-augmented conditioningon nearby street-view images. However, this design introduces severalchallenges, including temporal misalignment between retrieved referencesand the dynamic target scene, limited trajectory diversity and data spar-sity from vehicle-mounted captures at sparse intervals. We address thesechallenges through cross-temporal pairing, a large-scale synthetic datasetenabling diverse camera trajectories, and a view interpolation pipelinethat synthesizes coherent training videos from sparse street-view images.We further introduce a Virtual Lookahead Sink to stabilize long-horizongeneration by continuously re-grounding each chunk to a retrieved im-age at a future location. We evaluate SWM against recent video worldmodels across three cities: Seoul, Busan, and Ann Arbor. SWM out-performs existing methods in generating spatially faithful, temporallyconsistent, long-horizon videos grounded in actual urban environmentsover trajectories reaching hundreds of meters, while supporting diversecamera movements and text-prompted scenario variations.