Real-time spatial audio generation for interactive video is now feasible with streaming diffusion models, enabling world models to produce complete audiovisual experiences rather than silent videos.
WorldSonus adds realistic spatial sound to AI-generated video environments in real-time. It uses a streaming diffusion model to generate audio that matches video content, responds to text instructions mid-generation, and creates stereo sound that aligns with scene geometry and camera movement.