Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
Abstract
As audio-visual generative models evolve into world simula-tors, cross-modal synchronization stands as a critical proxy for assessingthe consistency of world dynamics and causality in generated content.However, existing evaluation metrics presume structural correctness, re-ducing synchronization to mere temporal alignment. Consequently, theyfail on generative outputs, especially when exhibiting structural hallu-cinations and asymmetric cross-modal relations, which currently man-date expert human annotation to assess synchronization. Thisdependency introduces a critical paradox: human evaluators rely on rel-ative, reference-dependent comparisons, whereas automated metrics re-quire reference-free, absolute scalars. We resolve this paradox by propos-ing a framework that distills relative human perception into a continu-ous, globally consistent metric. First, we introduce SynthSync, a datasetof generative failures ranked via pairwise human annotations. Second,we adapt the Omni-LLM equipped with a continuous latent projectionto translate relative human rankings into continuous absolute values.Third, we propose Real-Valued Group Relative Policy Optimization (R-GRPO) to internalize the global causal structure of synchronization vialistwise score distributions. Empirically, our metric achieves state-of-the-art human preference alignment. We leverage this estimator to establisha standardized benchmark, advancing AV-Gen assessment from low-levelsignal correlation to visually grounded causality.