2 citations · 3 across the 5 of their papers we have counts for
6 papers
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Zhilong Zheng, Letian Tao, Yang Guan +7
Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tu…
On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration
Yujie Yang, Zhilong Zheng, Shengbo Eben Li
Ensuring the safety of environmental exploration is a critical problem in reinforcement learning (RL). While limiting exploration to a feasible zone has become widely accepted as a…
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
Shiqi Liu, Zeyu He, Guojian Zhan +10
Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regu…
The Feasibility Theory of Constrained Reinforcement Learning: A Tutorial Study
Yujie Yang, Zhilong Zheng, Masayoshi Tomizuka +2
Satisfying safety constraints is a priority concern when solving optimal control problems (OCPs). Due to the existence of infeasibility phenomenon, where a constraint-satisfying so…
On the Stability of Datatic Control Systems
Yujie Yang, Zhilong Zheng, Shengbo Eben Li
The development of feedback controllers is undergoing a paradigm shift from (model-driven) control to (data-driven) control. Stability, as a f…
Feasible Policy Iteration for Safe Reinforcement Learning
Yujie Yang, Zhilong Zheng, Shengbo Eben Li +4
Safety is the priority concern when applying reinforcement learning (RL) algorithms to real-world control problems. While policy iteration provides a fundamental algorithm for stan…