1 citations · 1 across the 11 of their papers we have counts for
14 papers
From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation
Xiangyu Shi, Ruoxi Yang, Wei Tao +3
Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation p…
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
Qingyun Liu, Jiwen Zhang, Jingyi Hu +2
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored.…
LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination
Taishan Li, Jiwen Zhang, Siyuan Wang +2
Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible. This a…
Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation
Shijun Wan, Xuehai Wu, Jiwen Zhang +2
Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are la…
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
Jiwen Zhang, Xiangyu Shi, Siyuan Wang +3
Vision-and-Language Navigation (VLN) has recently benefited from Multimodal Large Language Models (MLLMs), enabling zero-shot navigation. While recent exploration-based zero-shot m…
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
Libo Sun, Jiwen Zhang, Siyuan Wang +1
Mobile GUI agents powered by large foundation models enable autonomous task execution, but frequent updates altering UI appearance and reorganizing workflows cause agents trained o…