Map2World: Segment Map Conditioned Text to 3D World Generation
Abstract
High-quality world-scale 3D data is scarce, making completeand generalizable 3D world generation difficult. Existing approaches typi-cally rely on either generated images/videos or domain-specific 3D worlddatasets, resulting in worlds that are often incomplete or restricted tonarrow domains. In this paper, we present Map2World, a map-conditionedframework that converts semantic layouts into complete and generalizablelarge-scale 3D worlds. Given a segment map with per-region text prompts,Map2World uses the map as a controllable world-level interface and lever-ages TRELLIS, a powerful pretrained 3D asset generator, as a general3D prior. To scale asset-level priors to world-level synthesis, Map2Worldenables arbitrary spatial expansion via MultiDiffusion, controls globalscale through initial noise optimization, and enhances local geometry andappearance with a dedicated enhancer. Experiments demonstrate thatMap2World generates complete, layout-controllable, scale-consistent, anddetailed 3D worlds across diverse layouts and semantics.