3 papers
cs.CL2026
Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation
Shijun Wan, Xuehai Wu, Jiwen Zhang +2
Social interaction depends on both language and visible social signals, such as facial expressions, posture, gaze, and emotional shifts. Yet existing social-agent benchmarks are la…
cs.AI2026
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
Jingcong Liang, Shijun Wan, Xuehai Wu +5
Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints.…
cs.CL2024
Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data
Xue Wu, Kostas Tsioutsiouliklis
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation. However, they often struggle with complex reasoning tasks a…