3 papers
cs.LG2026
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning
Abdelghani Ghanem, Mounir Ghogho
Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, i…
cs.LG2026
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
Abdelghani Ghanem, Mounir Ghogho
Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture complex, multi-modal behaviors. While Q-…
cs.LG2023
Multi-Objective Decision Transformers for Offline Reinforcement Learning
Abdelghani Ghanem, Philippe Ciblat, Mounir Ghogho
Offline Reinforcement Learning (RL) is structured to derive policies from static trajectory data without requiring real-time environment interactions. Recent studies have shown the…