1 paper
Qingyun Liu, Jiwen Zhang, Jingyi Hu +2
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored.…