Agents can now learn from and generate high-quality videos by working with structured knowledge graph representations instead of raw pixels, improving video generation quality by 20.7% over existing methods.
AVA-Encoder learns video representations as knowledge graphs that agents can reason about and edit. It converts videos into structured text and asset layers, then reconstructs videos from these representations. A natural-language feedback loop optimizes the encoding, enabling agents to work with cinematic-quality videos more effectively.