Generalizable Neural Reconstruction of High-Fidelity Surfaces via Sparse Volumetric Representations
Abstract
Neural implicit representations have recently achieved im-pressive results in novel view synthesis and multi-view 3D reconstruc-tion, yet both NeRF- and Gaussian Splatting-based methods require per-scene optimization, which makes them inefficient. Generalizable NeuralSurface Reconstruction (GNSR) methods have been proposed to removethis need by learning feature representations directly predicted from in-put images. However, their typical reliance on dense feature volumesseverely limits achievable resolution and fidelity due to prohibitive mem-ory costs. We introduce Sparse Volumetric Reconstruction (SVRecon),a new GNSR framework that unlocks high-resolution, memory-efficientreconstruction through learned occupancy-driven sparsity, in a more ef-fective way than earlier approaches to introducing sparsity in GNSRs.Our approach uses a nested two-stage architecture: (1) an occupancyprediction network that identifies surface-containing voxels, and (2) ahigh-resolution sparse volume rendering framework defined only withinthese occupied regions, together with specialized sparsified algorithmsfor ray sampling, feature aggregation, and querying. This design enablesfine-grained surface reconstruction while avoiding the heavy memoryfootprint of dense grids. SVRecon operates at resolutions up to 5123on standard 32GB hardware—substantially higher than prior generaliz-able methods—and delivers smoother and more precise reconstructionsacross diverse datasets, particularly in sparse-view settings.