Small language models can efficiently run agent cognition (thinking and memory) on edge devices like Jetson boards, enabling virtual agents to operate independently in real-time without cloud latency.
This paper explores how small language models (SLMs) running on edge devices can power the cognitive processes of virtual agents in immersive worlds.