2 papers
cs.LG2026
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning
Abdelghani Ghanem, Mounir Ghogho
Multi-step returns accelerate reward propagation in off-policy reinforcement learning, but couple the evaluation of each decision to the suboptimal logged actions that follow it, i…
cs.LG2026
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
Abdelghani Ghanem, Mounir Ghogho
Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture complex, multi-modal behaviors. While Q-…