Self-reflection (validating and correcting actions at runtime) is more important than memory for making language agents reliable in interactive environments, but combining both works best.
This paper improves language agents for interactive tasks by adding two cognitive modules to SwiftSage: a memory system that stores and retrieves important experiences, and a self-reflection system that validates actions and fixes mistakes. Testing on ScienceWorld shows the full system performs best, with self-reflection being the most critical component for handling long-horizon tasks.