VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
Abstract
Multi-expert models have become the dominant paradigmfor long-tailed learning, largely attributed to their presumed ability tobenefit from expert diversity. However, we revisit this central assump-tion and reveal that diversity induced by logit adjustment or explicitregularizers does not guarantee better ensemble accuracy. Our work sug-gests that multi-expert models benefit more from variance reduction thandiversity maximization. We introduce VICAL, a VIcinal ConsistencyALignment framework that improves long-tailed recognition not by en-forcing expert diversity, but by reducing prediction variance. Specifically,our approach comprises two key components: Self-Consistency Learn-ing and Deep Ensemble Distillation. Self-Consistency Learning discour-ages reliance on unstable high-frequency information, smoothing the lo-cal loss landscape and mitigating overfitting, especially for tail classes.Deep Ensemble Distillation promotes cross-expert low-frequency seman-tic agreement using a low-resolution view, thereby sidestepping opti-mization conflicts with established knowledge. Extensive experiments onCIFAR-LT, ImageNet-LT, and iNaturalist 2018 show that VICAL consis-tently outperforms state-of-the-art methods, validating the effectivenessof our consistency-driven design.