2 papers
cs.CV2026
Memory-Conditioned Tool Calling for Camera-First Visual Agents
Xiaofan Wu, Xi Zeng, Miaoxia Chen +5
Recognition tells an agent what is in an image; personal memory affects what is worth looking up next. In a camera-first setting the user can send only an image, so the agent must…
cs.RO2026
From World Models to World Action Models: A Concise Tutorial for Robotics
Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should…