2 papers
cs.CL2026
HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning
Yucan Guo, Xiaohan Wang, Miao Su +8
Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has be…
cs.CL2026
When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents
Zihan Lin, Zhenyu Chen, Jiawen Wei +6
Extracting skills from past successes is critical for the efficient evolution of Large Language Model (LLM) agents. Prevailing agent self-evolution paradigms typically rely on a co…