Using structured graphs to represent object relationships provides a more scalable and precise way to control video generation than trajectory drawing or text prompts, especially for complex multi-object scenes.
GraphVid enables precise control over multi-object interactions in video generation by using structured interaction graphs instead of text or pixel-level motion inputs. The method outperforms existing motion-control approaches while using less training data, and the authors release GraphVid-Bench, a new dataset with relational annotations for interaction-aware video generation.