Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
Abstract
Enabling embodied agents to imagine future states is essen-tial for robust and generalizable visual navigation. Yet, state-of-the-artsystems typically rely on modular designs that decouple navigation plan-ning from visual world modeling, which often induces state–action mis-alignment and weak adaptability in novel or dynamic scenarios. We pro-pose UniWM, a unified, memory-augmented world model that integratesegocentric visual foresight and planning within a single multimodal au-toregressive backbone. UniWM explicitly grounds action selection in vi-sually imagined outcomes, tightly aligning prediction with control. Mean-while, a hierarchical memory mechanism fuses short-term perceptual cueswith longer-term trajectory context, supporting stable and coherent rea-soning over extended horizons. Extensive experiments on four challeng-ing benchmarks (Go Stanford, ReCon, SCAND, HuRoN) and the 1XHumanoid Dataset show that UniWM improves navigation success ratesby up to 30%, substantially reduces trajectory errors against strong base-lines, generalizes zero-shot to the unseen TartanDrive dataset, and scalesnaturally to high-dimensional humanoid navigation. These results po-sition UniWM as a principled step toward unified, imagination-drivenembodied navigation. All materials are committed to be open-sourced.