GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
Abstract
The ex001Ecient spatial allocation of primitives serves as thefoundation of 3D Gaussian Splatting, as it directly dictates the synergybetween representation compactness, reconstruction speed, and render-ing x001Cdelity. Previous solutions, whether based on iterative optimizationor feed-forward inference, sux001Ber from signix001Ccant trade-ox001Bs between thesegoals, mainly due to the reliance on local, heuristic-driven allocationstrategies that lack global scene awareness. Specix001Ccally, current feed-forward methods are largely pixel-aligned or primitive-aligned. By un-projecting pixels into dense, view-aligned primitives, they bake redun-dancy into the 3D asset. As more input views are added, the represen-tation size increases and global consistency becomes fragile. To this end,we introduce GlobalSplat, a framework built on the principle of alignx001Crst, decode later. Our approach learns a compact, global, latent scenerepresentation that encodes multi-view input and resolves cross-view cor-respondences before decoding any explicit 3D geometry. Crucially, thisformulation enables compact, globally consistent reconstructions withoutrelying on pretrained pixel-prediction backbones or reusing latent fea-tures from dense baselines. Utilizing a coarse-to-x001Cne training curriculumthat gradually increases decoded capacity, GlobalSplat natively preventsrepresentation bloat. On RealEstate10K and ACID, our model achievescompetitive novel-view synthesis performance while utilizing as few as16K Gaussians, signix001Ccantly less than required by dense pipelines, ob-taining a light 4MB footprint. Further, GlobalSplat enables signix001Ccantlyfaster inference than the baselines, operating under 78 milliseconds in asingle forward pass.