2 citations · 2 across the 4 of their papers we have counts for
4 papers
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning
Dayang Liang, Lang Feng, Bo An +1
Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantag…
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
Dayang Liang, Liyuan He, Xuan Feng +3
Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants…
Episodic Reinforcement Learning with Expanded State-reward Space
Dayang Liang, Yaru Zhang, Yunlong Liu
Empowered by deep neural networks, deep reinforcement learning (DRL) has demonstrated tremendous empirical successes in various domains, including games, health care, and autonomou…
Sequential Action-Induced Invariant Representation for Reinforcement Learning
Dayang Liang, Qihang Chen, Yunlong Liu
How to accurately learn task-relevant state representations from high-dimensional observations with visual distractions is a realistic and challenging problem in visual reinforceme…