V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Abstract
We present V-JEPA 2.1, a family of self-supervised models that learns dense, high-quality, and temporally consistent representations for visual scenes in both images and videos. V-JEPA 2.1 combines four key ingredients: (i) a Dense Predictive Loss, a masking-based objective in which all tokens—visible context and masked tokens alike— contribute to the training loss, encouraging explicit spatial and temporal grounding; (ii) Deep Self-Supervision, which applies the selfsupervised objective hierarchically at multiple intermediate encoder layers to improve representation quality; (iii) Multi-Modal Tokenizers that support unified training over images and videos; and (iv) effective model and data scaling. Empirically, V-JEPA 2.1 achieves state-ofthe-art results on a range of benchmarks: 7.71 mAP on Ego4D for shortterm object-interaction anticipation, 40.8 Recall@5 on EPIC-Kitchens for high-level action anticipation, and a 20% improvement in real-robot grasping success rate over VJEPA-2 AC. The model also achieves stateof-art performance in robotic navigation (5.687 ATE on Tartan Drive), depth estimation (0.307 RMSE on NYUv2 with a linear probe), and global recognition (77.7% on Something-Something-V2).