Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
Abstract
Decades of orbital missions have produced multi-modal re-mote sensing data for the Moon, spanning optical imagery, spectroscopy,thermal emission, radar, gravity, and elemental composition. Yet thesedatasets remain fragmented across archives, and no benchmark exists forevaluating machine learning on lunar data. We introduce Moonstone, thefirst multi-modal foundation model benchmark for lunar remote sensing.Our contributions are: (1) a 28-channel, 128 pixels-per-degree (∼237 m)global lunar pretraining dataset from seven instrument families acrossfive missions, (2) MG-MAE, a modality-grouped masked autoencoderwith per-group convolutional tokenizers, a shared Vision Transformerencoder, attention masking for missing modalities, coverage-adaptivemasking for heterogeneous spatial coverage, and spectral continuity reg-ularization for physically plausible reconstructions, and (3) a benchmarkof six downstream tasks covering classification, regression, and segmen-tation. MG-MAE pretrained features outperform scratch baselines onall tasks and surpass both ImageNet-pretrained and vanilla MAE base-lines by large margins. We release the pretraining dataset, code, and thebenchmark suite.3