TriNLOS: Triplane Representations for Neural Non-Line-of-Sight Imaging
Abstract
Non-line-of-sight (NLOS) imaging aims to reconstruct hid-den scenes from time-resolved light transport, enabling vision around cor-ners. While classical physics-based methods provide principled inversion,learning-based approaches often rely on dense 3D backbones with cubiccomputational complexity. We propose a hybrid deep learning frame-work that combines a physics-guided initialization with a triplane-basedbackbone for high-fidelity volumetric reconstruction. The initialization isprovided by a learnable Enhanced Light-Cone Transform (ELCT), whichproduces a stable physics-consistent coarse volume, while the learnedbackbone replaces expensive O(N 3 ) 3D processing with scalable O(N 2 )triplane feature extraction. ConvNeXt-style residual convolutions, Restormerattention, and axis-aware cross-attention jointly refine structure and re-cover missing geometry. Experiments on synthetic and real NLOS datademonstrate improved reconstruction fidelity compared to representativephysics-based and learning-based baselines.