activity
20202026
most citedOptimization Landscape of Policy Gradient Methods for Discrete-time Static Output Feedback

10 citations · 35 across the 26 of their papers we have counts for

collaborators

33 papers

cs.RO2026

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion

Guanchen Lu, Yajuan Dun, Yi Zhou +4

Scalable reinforcement learning has popularized high-throughput sampling architectures, which significantly compresses the training time for off-policy methods in robotic locomotio…

eess.SY2026

On the Optimization Landscape of Observer-based Dynamic Linear Quadratic Control

Jingliang Duan, Jie Li, Yinsong Ma +5

Understanding the optimization landscape of linear quadratic regulation (LQR) problems is fundamental to the design of efficient reinforcement learning solutions. Recent work has m…

cs.CL2026

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

Shiqi Liu, Zeyu He, Guojian Zhan +10

Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regu…

cs.AI2025

Off-policy Reinforcement Learning with Model-based Exploration Augmentation

Likun Wang, Xiangteng Zhang, Yinuo Wang +5

Exploration is fundamental to reinforcement learning (RL), as it determines how effectively an agent discovers and exploits the underlying structure of its environment to achieve o…

cs.LG2025

Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL

Guojian Zhan, Likun Wang, Pengcheng Wang +4

Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further pe…

cs.LG2025

Distributional Soft Actor-Critic with Diffusion Policy

Tong Liu, Yinuo Wang, Xujie Song +6

Reinforcement learning has been proven to be highly effective in handling complex control tasks. Traditional methods typically use unimodal distributions, such as Gaussian distribu…