collaborators

14 papers

cs.LG2026

Complementary RL: Towards Efficient Experience-Driven Agent Learning

Dilxat Muhtar, Jiashun Liu, Wei Gao +8

Reinforcement Learning (RL) has emerged as a powerful paradigm for training LLM-based agents, yet remains limited by low sample efficiency, stemming not only from sparse outcome fe…

cs.LG2026

Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy

Jiashun Liu, Runze Liu, Xu Wan +3

Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or even training collapses. Recent…

cs.CL2026

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization

Yang Li, Zhichen Dong, Yuhan Sun +9

The reasoning pattern of Large language models (LLMs) remains opaque, and reinforcement learning (RL) typically applies uniform credit across an entire generation, blurring the dis…

cs.LG2026

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

Runze Liu, Jiashun Liu, Xu Wan +2

Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT is expected to provide a usefu…

cs.LG2026

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions

Bingxu Liu, Jiashun Liu, Johan Obando-Ceron +5

While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stat…

cs.AI2026

Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem

Weixun Wang, XiaoXiao Xu, Wanhe An +86

Agentic crafting requires LLMs to operate in real-world environments over multiple turns by taking actions, observing outcomes, and iteratively refining artifacts. Despite its impo…