MuCHeR: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking
Abstract
Most advances in human mesh recovery (HMR) have focusedon pelvis-centered recovery, overlooking metric 3D localization and detec-tion accuracy in the camera coordinate system – two key factors for real-world applications such as human–robot interaction and social scene un-derstanding. Current evaluation protocols often ignore these aspects, em-phasizing per-person, root-centered recovery rather than camera-spaceperception. As a result, existing approaches rely on fixed camera as-sumptions or handcrafted post-processing, limiting their robustness andpractical deployment. We introduce Multi-HMR 2, a simple yet ro-bust DETR-based framework for Multi-person camera-centric Humandetection, Mesh Recovery, and tracking. Multi-HMR 2 predicts a scene-consistent camera together with human meshes, enabling metric 3D lo-calization without ground-truth intrinsics. Moreover, by distilling image-based memory features from SAM2, Multi-HMR 2 extends to tracking,achieving consistent identity association without video supervision. De-spite its conceptual simplicity – no handcrafted components, no videoinput, and no ground-truth cameras – Multi-HMR 2 achieves state-of-the-art pelvis-centered performance while substantially improving detec-tion accuracy and metric 3D localization.