Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
Shichao Ma, Zhiyuan Ma, Ming Yang +8
Multi-turn tool-integrated reasoning enables Large Language Models (LLMs) to solve complex tasks through iterative information retrieval. However, current reinforcement learning (R…
cs.AI2025
DART: Difficulty-Adaptive Reasoning Truncation for Efficient Large Language Models
Ruofan Zhang, Bin Xia, Zhen Cheng +4
Adaptive reasoning is essential for aligning the computational effort of large language models (LLMs) with the intrinsic difficulty of problems. Current chain-of-thought methods bo…