SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for RGB+Thermal 3D Reconstruction
Abstract
Foundational feed-forward visual geometry models enableaccurate and efficient camera pose estimation and scene reconstruction bylearning strong scene priors from massive RGB datasets. However, theireffectiveness drops when applied to mixed sensing modalities, such asRGB-thermal (RGB-T) images. We observe that while a visual geometrygrounded transformer pretrained on RGB data generalizes well to thermal-only reconstruction, it struggles to align RGB and thermal modalitieswhen processed jointly. To address this, we propose SEAR, a simple yet ef-ficient fine-tuning strategy that adapts a pretrained geometry transformerto multimodal RGB-T inputs. Despite being trained on a relatively smallRGB-T dataset, our approach significantly outperforms state-of-the-artmethods for 3D reconstruction and camera pose estimation, achievingsignificant improvements over all metrics and delivering higher detailand consistency between modalities with negligible overhead in inferencetime compared to the original pretrained model. Notably, SEAR enablesreliable multimodal pose estimation and reconstruction even under chal-lenging conditions, such as low lighting and dense smoke. We validateour architecture through extensive ablation studies and demonstratehow the model aligns both modalities. Additionally, we introduce a newdataset featuring RGB and thermal sequences captured at different times,viewpoints, and illumination conditions, providing a robust benchmarkfor future work in multimodal 3D scene reconstruction. Code and modelsare publicly available at https://doi.org/10.5281/ZENODO.21077295.