3 papers
cs.AI2026
NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities
Haonan Jiang, Guojian Zhan, Jiancong Xie +6
Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmar…
cs.CL2026
Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation
Shijun Wan, Xuehai Wu, Jiwen Zhang +2
Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are la…
cs.AI2025
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
Jingcong Liang, Shijun Wan, Xuehai Wu +5
Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints.…