3 papers
cs.LG2026
Efficient Preference Poisoning Attack on Offline RLHF
Chenye Yang, Weiyu Xu, Lifeng Lai
Offline Reinforcement Learning from Human Feedback (RLHF) pipelines such as Direct Preference Optimization (DPO) train on a pre-collected preference dataset, which makes them vulne…
cs.LG2025
Learn to Change the World: Multi-level Reinforcement Learning with Model-Changing Actions
Ziqing Lu, Babak Hassibi, Lifeng Lai +1
Reinforcement learning usually assumes a given or sometimes even fixed environment in which an agent seeks an optimal policy to maximize its long-term discounted reward. In contras…
cs.LG2025
Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach
Ziqing Lu, Lifeng Lai, Weiyu Xu
Reinforcement learning (RL) for the Markov Decision Process (MDP) has emerged in many security-related applications, such as autonomous driving, financial decisions, and drone/robo…