15 papers
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
Yi Ding, Ziliang Qiu, Bolian Li +1
Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs). However, existing reinforcement learning (RL) methods struggle to learn it, as…
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
Bolian Li, Yifan Wang, Yi Ding +3
Reinforcement learning (RL) has enabled complex reasoning abilities in large language models (LLMs). However, most RL algorithms suffer from performance saturation, preventing cont…
SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology
Yifan Wang, Bolian Li, David Cho +3
Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard to use in open-ended domains whe…
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
Anamika Lochab, Bolian Li, Ruqi Zhang
Reinforcement Learning with Verifiable Rewards (RLVR) has achieved substantial gains in single-attempt accuracy (Pass@1) on reasoning tasks, yet often suffers from reduced multi-sa…
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
Yifan Wang, Bolian Li, Junlin Wu +5
Real-world large language model deployments (e.g., conversational AI systems, code generation assistants) naturally generate abundant implicit user dissatisfaction (DSAT) signals,…
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
Zehong Wang, Fang Wu, Hongru Wang +8
Large language model (LLM)-based agents exhibit strong step-by-step reasoning capabilities over short horizons, yet often fail to sustain coherent behavior over long planning horiz…