4 papers
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
Chenhua Fan, Jiahui Zhu, Yuhang Zhang +1
Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies t…
An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints
Jiahui Zhu, Kihyun Yu, Dabeen Lee +2
Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn op…
ONER: Online Experience Replay for Incremental Anomaly Detection
Yizhou Jin, Jiahui Zhu, Guodong Wang +5
Incremental anomaly detection aims to sequentially identify defects in industrial product lines but suffers from catastrophic forgetting, primarily due to knowledge overwriting dur…
A Survey on Data Synthesis and Augmentation for Large Language Models
Ke Wang, Jiahui Zhu, Minjie Ren +8
The success of Large Language Models (LLMs) is inherently linked to the availability of vast, diverse, and high-quality data for training and evaluation. However, the growth rate o…