Contextualizing visual instructions to match a user's real workspace—through AI-generated images and videos—improves task performance and trust compared to generic pre-authored tutorials.
This paper presents a system that generates live visual instructions tailored to a user's specific workspace and task progress, rather than showing pre-recorded tutorials. Using AR and generative AI, it creates goal images and demo videos that match the user's actual environment, helping them complete physical tasks more accurately and confidently.