collaborators

13 papers

cs.LG2026

Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control

Qi Zhao, Guozheng Ma, Yilun Kong +9

Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many…

cs.RO2026

ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning

Yilun Kong, Yunpeng Qing, Guozheng Ma +4

The paper proposes ExToken, a framework that conditions vision‑language‑action policies on discrete behavioral tokens derived from offline demonstrations to promote diverse, struct…

cs.LG2026

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

Guozheng Ma, Lu Li, Zilin Wang +2

Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradi…

cs.AI2026

Language-based Trial and Error Falls Behind in the Era of Experience

Haoyu Wang, Guozheng Ma, Shugang Cui +7

While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.g., symbolic or spatial tasks) remains limite…

cs.LG2026

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning

Junjie Zhang, Guozheng Ma, Shunyu Liu +5

Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs). However, existing step-level…

cs.CL2026

A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models

Junjie Zhang, Guozheng Ma, Shunyu Liu +6

Reinforcement Learning with Verifiable Rewards~(RLVR) has emerged as a powerful learn-to-reason paradigm for large reasoning models to tackle complex tasks. However, the current RL…