Open-Vocabulary Long Term Action Anticipation
Abstract
Action anticipation, i.e., predicting future actions from past video observations, is fundamental to intelligent systems that assist humans. Despite progress in model architectures, current evaluation practices exhibit critical limitations: methods evaluate exclusively in closedset settings where the training and test action vocabularies are identical, thereby preventing an understanding of generalization to novel action classes encountered in real-world deployment. We therefore introduce the first open-vocabulary evaluation framework for action anticipation, where models are trained on one egocentric dataset and tested on entirely different egocentric datasets with novel action vocabularies. Since our thorough evaluation shows that adapting existing approaches to this task is insufficient, we propose a novel approach that employs horizonspecific learnable queries and a lightweight text encoder adaptation for open-vocabulary long-term action anticipation. It substantially outperforms other approaches that we have adapted to this task.