CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection
Abstract
Few-shot anomaly detection remains challenging due to lim-ited training data. Multi-modal anomaly detection (MAD) offers a viablesolution, leveraging 3D geometric cues to enrich 2D RGB representationsand compensate for this scarcity. However, existing MAD methods applyspatially uniform feature processing, conflating stable macroscopic struc-tures with high-frequency localized defect signals, exacerbating cross-modal misalignment and inflating false-positive rates. To overcome this,we present CMDS-AD, a Cross-Modal Dual-Stream Anomaly Detectionframework. A LoRA-guided diffusion model generates diverse RGB sam-ples to mitigate extreme data scarcity. For 3D normal augmentation, weemploy a pre-trained diffusion model as a normal estimator. Crucially,this estimator inherently acts as a non-linear low-pass filter, directly ex-tracting low-frequency normal representations from RGB inputs. Thisestablishes an auxiliary estimated stream of purely low-frequency infor-mation, anchoring robust structural templates and assisting the uncom-pressed real stream, containing coupled high- and low-frequency compo-nents, to precisely isolate micro-defects. A Coordinate-Aware Hierarchi-cal Feature Mapper adaptively aligns cross-modal semantics, while a mul-tiplicative scoring mechanism filters modality-specific noise. Under theextreme 1-shot setting, CMDS-AD achieves absolute performance gainsof 5.7% (I-AUROC) and 2.0% (AUPRO) on MVTec 3D-AD, along-side 7.7% and 5.6% improvements on EyeCandies, establishing a newstate-of-the-art. Code is available at Junhaocai27/CMDS-AD.