1 paper · 1 filter
Yutong Wu, Jie Zhang, Yiming Li +4
Vision Language Model (VLM)-based agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent sy…