13 papers
Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control
Qi Zhao, Guozheng Ma, Yilun Kong +9
Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many…
ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning
Yilun Kong, Yunpeng Qing, Guozheng Ma +4
The paper proposes ExToken, a framework that conditions vision‑language‑action policies on discrete behavioral tokens derived from offline demonstrations to promote diverse, struct…
Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors
Guozheng Ma, Lu Li, Zilin Wang +2
Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradi…
Language-based Trial and Error Falls Behind in the Era of Experience
Haoyu Wang, Guozheng Ma, Shugang Cui +7
While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.g., symbolic or spatial tasks) remains limite…
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning
Junjie Zhang, Guozheng Ma, Shunyu Liu +5
Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs). However, existing step-level…
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
Junjie Zhang, Guozheng Ma, Shunyu Liu +6
Reinforcement Learning with Verifiable Rewards~(RLVR) has emerged as a powerful learn-to-reason paradigm for large reasoning models to tackle complex tasks. However, the current RL…