collaborators

5 papers

cs.CL2026

PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models

Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang +6

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is large…

cs.CV2025

TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning

Ming Li, Jike Zhong, Shitian Zhao +6

The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as t…

cs.AI2025

SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language Models

Zhongjian Miao, Hao Fu, Chen Wei

We introduce SPAN, a cross-calendar temporal reasoning benchmark, which requires LLMs to perform intra-calendar temporal reasoning and inter-calendar temporal conversion. SPAN feat…

cs.CL2025

PyVision: Agentic Vision with Dynamic Tooling

Shitian Zhao, Haoquan Zhang, Shaoheng Lin +4

LLMs are increasingly deployed as agents, systems capable of planning, reasoning, and dynamically calling external tools. However, in visual reasoning, prior approaches largely rem…

cs.CL2025

Exploring Large Language Models for Word Games:Who is the Spy?

Chentian Wei, Jiewei Chen, Jinzhu Xu

Word games hold significant research value for natural language processing (NLP), game theory, and related fields due to their rule-based and situational nature. This study explore…