activity
20232026
most citedFeasible Policy Iteration for Safe Reinforcement Learning

2 citations · 3 across the 5 of their papers we have counts for

collaborators

6 papers

cs.LG2026

Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification

Zhilong Zheng, Letian Tao, Yang Guan +7

Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tu…

cs.LG2026

On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration

Yujie Yang, Zhilong Zheng, Shengbo Eben Li

Ensuring the safety of environmental exploration is a critical problem in reinforcement learning (RL). While limiting exploration to a feasible zone has become widely accepted as a…

cs.CL2026

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

Shiqi Liu, Zeyu He, Guojian Zhan +10

Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regu…

eess.SY2024

The Feasibility Theory of Constrained Reinforcement Learning: A Tutorial Study

Yujie Yang, Zhilong Zheng, Masayoshi Tomizuka +2

Satisfying safety constraints is a priority concern when solving optimal control problems (OCPs). Due to the existence of infeasibility phenomenon, where a constraint-satisfying so…

eess.SY2024★ 1 cited

On the Stability of Datatic Control Systems

Yujie Yang, Zhilong Zheng, Shengbo Eben Li

The development of feedback controllers is undergoing a paradigm shift from (model-driven) control to (data-driven) control. Stability, as a f…

cs.LG2023★ 2 cited

Feasible Policy Iteration for Safe Reinforcement Learning

Yujie Yang, Zhilong Zheng, Shengbo Eben Li +4

Safety is the priority concern when applying reinforcement learning (RL) algorithms to real-world control problems. While policy iteration provides a fundamental algorithm for stan…