1 paper
Zongxin Yang, Guikun Chen, Xiaodi Li +2
Recent LLM-driven visual agents mainly focus on solving image-based tasks, which limits their ability to understand dynamic scenes, making it far from real-life applications like g…