A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
Abstract
Trained policies for real-world robotics rely on computer vi-sion components, typically in the form of pre-trained visual encoders.These encoders are an essential component and it has been shown thattheir power does not emerge from training on robotics downstream lossesalone. Pre-training with auxiliary losses in the form of computer-visionpre-text tasks is a defining factor and heavily conditions agent perfor-mance in robotics tasks. In this unprecedented large-scale study, we ran966 navigation episodes of static point goal navigation in a real-worldbuilding for 24km and asked which components really matter for thecomputer vision aspects of robotics: we evaluate state-of-the art visualencoders in realistic conditions. We explore the usefulness of heteroge-neous multi-teacher distillation leading to encoders with multiple dif-ferent and complementary skills. We investigate how much informationfrom these encoders is necessary for robotics by bottlenecking them in aprincipled and “spatially useful” way and we show that this leads to theemergence of interpretable features linked to affordances. We also arguethat training policies on RGB data alone does not lead to an optimalusage of visual features and show this by finetuning policies pre-trainedon privileged information. All in all, we paint a more complete picture ofwhat aspects of computer vision are relevant for real-world navigation.