4 papers
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
Kongcheng Zhang, Qi Yao, Shunyu Liu +7
Reinforcement Learning (RL) has shown promise for aligning Large Language Models (LLMs) to follow instructions with various constraints. Despite the encouraging results, RL improve…
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Kongcheng Zhang, Qi Yao, Shunyu Liu +5
Recent advances of Reinforcement Learning (RL) have highlighted its potential in complex reasoning tasks, yet effective training often relies on external supervision, which limits…
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
Qingmao Yao, Zhichao Lei, Tianyuan Chen +5
Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the -value overestimation for out-of-distribution (OOD) actions. Existing methods address th…
Reasoning with Reinforced Functional Token Tuning
Kongcheng Zhang, Qi Yao, Baisheng Lai +5
In this work, we propose Reinforced Functional Token Tuning (RFTT), a novel reinforced fine-tuning framework that empowers Large Language Models (LLMs) with self-play learn-to-reas…