19 citations · 36 across the 17 of their papers we have counts for
10 papers · 1 filter
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
Shiqi Liu, Zeyu He, Letian Tao +9
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and op…
On the Identifiability of Controlled World Models
Xiangteng Zhang, Yang Guan, Bo Zhang +3
World model serves as a promising tool to infer environment dynamics under high-dimensional observations and candidate actions. Recently, LeCun's JEPA provides a compelling framewo…
FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
Bonan Wang, Letian Tao, Bin Shuai +7
Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates…
Model-based Chance-Constrained Reinforcement Learning via Separated Proportional-Integral Lagrangian
Baiyu Peng, Jingliang Duan, Jianyu Chen +6
Safety is essential for reinforcement learning (RL) applied in the real world. Adding chance constraints (or probabilistic constraints) is a suitable way to enhance RL safety under…
Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety
Haitong Ma, Yang Guan, Shegnbo Eben Li +3
The safety constraints commonly used by existing safe reinforcement learning (RL) methods are defined only on expectation of initial states, but allow each certain state to be unsa…
Integrated Decision and Control: Towards Interpretable and Computationally Efficient Driving Intelligence
Yang Guan, Yangang Ren, Qi Sun +5
Decision and control are core functionalities of high-level automated vehicles. Current mainstream methods, such as functionality decomposition and end-to-end reinforcement learnin…