VOCA: Visual Odometry with Codec Awareness
Abstract
Camera pose estimation from image streams is a critical com-ponent of spatial world models that integrate perception into planningand decision-making. Nearly all Visual Odometry (VO) and SimultaneousLocalization and Mapping (SLAM) systems have focused on datasetscontaining raw, uncompressed videos. Many working systems insteaduse ubiquitous hardware units to efficiently compress and decode videostreams, saving orders of magnitude in storage and bandwidth. However,this lossy compression introduces visual artifacts that hinder the per-formance of traditional tracking systems. We present VOCA, a causalstereo visual-odometry method that exploits codec information to im-prove tracking performance. We achieve state-of-the-art performance oncausal VO for relative trajectory error, efficiency, and absolute trajec-tory error on compressed streams. This work highlights the potential ofleveraging widely available video codec information for vision tasks.