collaborators

29 papers

cs.MA2026

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

Qingyun Liu, Jiwen Zhang, Jingyi Hu +2

Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored.…

cs.CY2026

Uncovering Salience-Driven Dynamics in Consumer Confidence with Generative Social Simulation

Yixu Huang, Yunlu Yin, Jiayu Lin +6

Consumer confidence is typically modeled as a persistent macroeconomic index, yet its movements arise from households that interpret economic information through heterogeneous cons…

cs.CV2026

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination

Taishan Li, Jiwen Zhang, Siyuan Wang +2

Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible. This a…

cs.CL2026

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

Shijun Wan, Xuehai Wu, Jiwen Zhang +2

Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are la…

cs.CL2026

HyLaT: Efficient Multi-Agent Communication via Hybrid Latent-Text Protocol

Xinyi Mou, Siyuan Wang, Zejun Li +2

Communication protocol design is a central challenge in large language model-based multi-agent systems. Existing single-channel approaches face an inherent communication trilemma:…

cs.AI2026

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning

Zejun Li, Yingxiu Zhao, Jiwen Zhang +6

Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be achieved in particular domains, they struggle to develop general r…