C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation
Abstract
LiDAR-based 3D object detection is essential for autonomousdriving systems. However, traditional Ego-only Perception (Eo-Perception)suffers from limited perspective and occlusions in a complex outdoor en-vironment, leading to performance bottlenecks. Recently, research onmulti-agent Collaborative Perception (Co-Perception) has demonstratedexcellent performance, but high communication costs and accumulatedpose error hinder its application. To address this, we explore a novelC2E (Co-Perception to Eo-Perception) paradigm through the Multi-to-Single (M2S) agent contrastive knowledge distillation framework. OurM2S framework first designs Multi-Level Feature Enhancement moduleto provide more stable features, and introduces Auxiliary Point CloudReconstruction and Multi-Teacher Contrastive Distillation mechanismsto mitigate domain gaps in point cloud and feature distributions withinthe C2E paradigm. Benefiting from this, our M2S can retain the ex-cellent performance of collaborative perception while effectively avoidingthe drawbacks, such as communication delays and positioning errors. Ex-tensive experiments on the V2XSet, V2V4Real and DAIR-V2X datasetsshow the effectiveness and generalizability of our M2S framework whencombined with the state-of-the-art CoSDH model and other excellent 3Ddetectors. Our M2S framework can deliver up to a 8.64% improvementin 3D mAP performance without introducing any communication costs.