JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Abstract
In this paper, we present JoVA, a streamlined frameworkthat unifies joint video-audio generation and editing. While existingmethods often rely on fragmented, task-specific architectures or complexfusion mechanisms, JoVA employs native joint representation learningfor direct video, audio, and text interaction in a dual-branch architec-ture. This design eliminates redundant alignment modules and effectivelyunifies diverse multimodal tasks within a single model. Furthermore, weutilize channel-wise conditioning for flexible image and video reference toavoid massive token expansion, alongside a mouth-area loss to enhancelip alignment. To fully empower and systematically evaluate this frame-work, we construct a comprehensive training corpus encompassing video-audio generation and editing datasets, and introduce unified benchmarkstailored for these multimodal tasks. Extensive experiments demonstratethat JoVA achieves state-of-the-art performance across benchmarks, es-tablishing it as an extensible framework for versatile content creation.Project page: https://visual-ai.github.io/jova