works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.LG2026

FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving

Bonan Wang, Letian Tao, Bin Shuai +7

The paper introduces FAST, a synchronous parallel framework that improves sampling efficiency for deep reinforcement learning in autonomous driving by aligning parallel simulations…

cs.RO2026

FastDSAC: Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion

Guanchen Lu, Yajuan Dun, Yi Zhou +4

Scalable reinforcement learning has popularized high-throughput sampling architectures, which significantly compresses the training time for off-policy methods in robotic locomotio…

cs.CL2026

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

Shiqi Liu, Zeyu He, Guojian Zhan +10

Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regu…

cs.LG2026

Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation

Guojian Zhan, Letian Tao, Pengcheng Wang +6

Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling com…

cs.LG2026

Real-Time Generative Policy via Langevin-Guided Flow Matching for Autonomous Driving

Tianze Zhu, Yinuo Wang, Wenjun Zou +6

Reinforcement learning (RL) is a fundamental methodology in autonomous driving systems, where generative policies exhibit considerable potential by leveraging their ability to mode…

cs.LG2024

Conformal Symplectic Optimization for Stable Reinforcement Learning

Yao Lyu, Xiangteng Zhang, Shengbo Eben Li +5

Training deep reinforcement learning (RL) agents necessitates overcoming the highly unstable nonconvex stochastic optimization inherent in the trial-and-error mechanism. To tackle…