VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) has emergedas a highly effective solution for novel view synthesis. Existing meth-ods predominantly rely on a pixel-aligned Gaussian prediction paradigm,where each 2D pixel is mapped to a 3D Gaussian. We rethink this widelyadopted formulation and identify several inherent limitations: it rendersthe reconstructed 3D models heavily dependent on the number of inputviews, leads to view-biased density distributions, and introduces align-ment errors, particularly when source views contain occlusions or lowtexture. To address these challenges, we introduce VolSplat, a new multi-view feed-forward paradigm that replaces pixel alignment with voxel-aligned Gaussians. By directly predicting Gaussians from a predicted3D voxel grid, it overcomes pixel alignment’s reliance on error-prone 2Dfeature matching, ensuring robust multi-view consistency. Furthermore,it enables adaptive control over density based on 3D scene complex-ity, yielding more faithful Gaussians, improved geometric consistency,and enhanced novel-view rendering quality. Experiments on widely usedbenchmarks demonstrate that VolSplat achieves state-of-the-art perfor-mance, while producing more plausible and view-consistent results.