Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation
Shijun Wan, Xuehai Wu, Jiwen Zhang +2
Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are la…
cs.CL2025
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
Xuanwen Ding, Chengjun Pan, Zejun Li +3
Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. T…
cs.CL2024
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Jiwen Zhang, Jihao Wu, Yihua Teng +5
Large language model (LLM) leads to a surge of autonomous GUI agents for smartphone, which completes a task triggered by natural language through predicting a sequence of actions o…