3 papers
cs.LG2026
CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning
Jiayi Guan, Tianle Zhang, Li Shen +8
Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks. This paradigm provides an effective mean…
cs.LG2025
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
Shangding Gu, Bilgehan Sel, Yuhao Ding +4
Ensuring the safety of Reinforcement Learning (RL) is crucial for its deployment in real-world applications. Nevertheless, managing the trade-off between reward and safety during e…
cs.CL2025
TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning
Shangding Gu, Alois Knoll, Ming Jin
The development of Large Language Models (LLMs) often confronts challenges stemming from the heavy reliance on human annotators in the reinforcement learning with human feedback (R…