12 papers
Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use
Zhixin Ma, Yutong Zhou, Yongqi Li +2
Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the "brain" of embodied AI, instructing robots to interact with the physical worl…
Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text
Yutong Bian, Dongjie Cheng, Heming Xia +2
Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves fr…
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
Qiancheng Xu, Yongqi Li, Fan Liu +3
Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools during reasoning. Existing TIR meth…
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
Heming Xia, Cunxiao Du, Rui Li +3
Large reasoning models (LRMs) have demonstrated remarkable proficiency in tackling complex tasks through step-by-step thinking. However, this lengthy reasoning process incurs subst…
Agent-as-a-Judge
Runyang You, Hongru Cai, Caiqi Zhang +5
LLM-as-a-Judge has revolutionized AI evaluation by leveraging large language models for scalable assessments. However, as evaluands become increasingly complex, specialized, and mu…
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
Jian Wang, Boyan Zhu, Chak Tou Leong +2
Large reasoning models (LRMs) have exhibited the capacity of enhancing reasoning performance via internal test-time scaling. Building upon this, a promising direction is to further…