4 papers · 1 filter
PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
Yu Li, Guangfeng Cai, Shengtian Yang +5
Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool pl…
Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning
Siyu Gong, Linan Yue, Weibo Gao +4
Tool-Integrated Reasoning (TIR) enables large language models (LLMs) to solve complex tasks by interacting with external tools, yet existing approaches depend on high-quality synth…
Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
Zongqian Wu, Baoduo Xu, Tianyu Li +3
Test-time scaling (TTS) has been shown to improve the performance of large language models (LLMs) by sampling and aggregating diverse reasoning paths. However, existing research ha…
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
Zongqian Wu, Tianyu Li, Baoduo Xu +4
Deep iterative chain-of-thought (CoT) reasoning enables LLMs to tackle complex tasks by progressively activating relevant pre-trained knowledge. However, it faces challenges in ens…