Adding a simple visual memory system that preserves and lets models retrieve past observations dramatically improves multimodal models' reasoning in interactive visual environments—achieving perfect performance on challenging puzzle games.
VISTA is a visual harness that enhances multimodal models' ability to solve complex interactive visual tasks by giving them long-horizon vision and lossless visual memory. The system lets models directly perceive environments, store past observations, and actively retrieve them while reasoning.