GraphVid: Interactive Graph-Controllable Video Generation
Abstract
Controllable video generation remains challenging due tothe di"culty of specifying precise multi-object interactions using textprompts or motion-control inputs that primarily constrain pixel move-ment. In practice, trajectory-based control often requires users to drawaccurate tracks for multiple objects, which scales poorly with scene com-plexity and becomes ambiguous under occlusion or overlap. To enableflexible yet precise multi-subject control, we introduce GraphVid, agraph-conditioned image-to-video generation model that enables inter-active control through structured interaction graphs. We further cu-rate GraphVid-Bench, a large-scale interaction-centric video datasetwith structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less trainingdata and fewer trainable parameters than prior motion-control methods,GraphVid delivers strong controllability and video quality. Comparedwith Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by37.6%, while improving PSNR (9.87→15.98) and SSIM (0.38→0.61). Ourresults highlight the potential of structured semantic interfaces as a pow-erful paradigm for controllable video generation.PLAN Lab https://plan-lab.github.io/graphvid