4 papers
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
Peiji Li, Linyang Li, Handa Sun +15
Large language models have demonstrated strong reasoning capabilities in complex tasks through tool integration, which is typically framed as a Markov Decision Process and optimize…
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
Yiliu Sun, Zicheng Zhao, Yang Wei +2
Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capability of Large Language Models (LLMs). Current RLVR approaches typically conduct tra…
CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate
Yiliu Sun, Zicheng Zhao, Sheng Wan +1
Nowadays, single Large Language Model (LLM) struggles with critical issues such as hallucination and inadequate reasoning abilities. To mitigate these issues, Multi-Agent Debate (M…
Fast-Slow-Thinking: Complex Task Solving with Large Language Models
Yiliu Sun, Yanfang Zhang, Zicheng Zhao +3
Nowadays, Large Language Models (LLMs) have been gradually employed to solve complex tasks. To face the challenge, task decomposition has become an effective way, which proposes to…