LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
Abstract
Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strictconstraints remains largely unverified. To address this, we introducea parametric safety benchmark to procedurally generate safety-criticalscenarios with comprehensive stochasticity. To overcome the scalabilitybottlenecks of human teleoperation, we develop a novel keypose-drivendata generation pipeline. Leveraging this infrastructure, we curate alarge-scale dataset of 19,664 strictly collision-free demonstrations withextensive domain randomization. We then conduct a systematic cross-paradigm evaluation of eight VLA and two embodied foundation mod-els. Our analysis reveals a critical generalization-safety tension: althoughhigh-diversity training fosters safer trajectories, task success remains fun-damentally bottlenecked by sub-optimal trajectory synthesis and seman-tic misalignment. By providing a scalable pipeline, a robust dataset, andprofound failure-mode insights, LIBERO-Safety establishes a crucialfoundation for developing safe and reliable VLA models.