Multimodal agents need explicit architectural constraints to use visual tools and external knowledge rather than defaulting to text search and internal memory; decoupled perception-exploration pipelines with staged tool unlocking significantly improve performance.
Video-DeepResearch extends multimodal AI agents to handle continuous video streams with web search integration. The system addresses two key problems: agents ignoring visual information in favor of text search, and relying on memorized knowledge instead of actually using tools.