Rethinking Cross-Spectral Image Generation via Shared-Specific Representation
Abstract
Cross-spectral image synthesis from visible to infrared facessignix001Ccant challenges due to the severe scarcity of infrared data in criti-cal scenarios. Existing methods primarily achieve style transfer throughpixel-level mappings, which often sux001Ber from semantic misalignment andunrealistic spectral characteristics. To address this problem, we rethinkcross-spectral generation as a decoupling and reconstruction process ofshared-specix001Cc features between modalities and propose SHASP, a novelgenerative network. Specix001Ccally, an encoding-decoupling module is de-signed at the front of the generator, which consists of a structure contentencoder and a spectral feature encoder to decouple the shared-specix001Ccfeatures. A decoding-reconstruction module is designed at the backendof the generator to perform feature combination and reconstruction. Wefurther impose semantic consistency constraints on shared features toensure precise cross-modal alignment. Extensive experiments on aerial,driving, and monitoring scenarios demonstrate that our method outper-forms state-of-the-art approaches. This work reveals that there are com-mon representations that can be mined at the semantic structure levelin cross-spectral data, and modality gaps can be accurately modeled viaa decoupling-reconstruction mechanism.