6 papers · 1 filter
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
Siyu Gong, Linan Yue, Weibo Gao +4
Tool-Integrated Reasoning (TIR) enables large language models (LLMs) to solve complex tasks by interacting with external tools, yet existing approaches depend on high-quality synth…
Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
Linan Yue, Yichao Du, Yizhi Wang +8
Recently, Large Reasoning Models (LRMs) have gradually become a research hotspot due to their outstanding performance in handling complex tasks. Among them, DeepSeek R1 has garnere…
IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response Theory
Wei Song, Zhenya Huang, Cheng Cheng +5
Large language models (LLMs) have demonstrated exceptional performance across a wide range of natural language tasks. However, selecting the optimal LLM to respond to a user query…
TestAgent: An Adaptive and Intelligent Expert for Human Assessment
Junhao Yu, Yan Zhuang, YuXuan Sun +5
Accurately assessing internal human states is key to understanding preferences, offering personalized services, and identifying challenges in real-world applications. Originating f…
CoderAgent: Simulating Student Behavior for Personalized Programming Learning with Large Language Models
Yi Zhan, Qi Liu, Weibo Gao +5
Personalized programming tutoring, such as exercise recommendation, can enhance learners' efficiency, motivation, and outcomes, which is increasingly important in modern digital ed…
A Survey of Models for Cognitive Diagnosis: New Developments and Future Directions
Fei Wang, Weibo Gao, Qi Liu +8
Cognitive diagnosis has been developed for decades as an effective measurement tool to evaluate human cognitive status such as ability level and knowledge mastery. It has been appl…