activity
20202026
most citedOpen RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning

3 citations · 6 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Drift Q-Learning

Anas Houssaini, Mohamad H. Danesh, Amin Abyaneh +3

Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies h…

cs.LG2026

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh +4

Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stoc…

cs.LG2025

Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation

Mohamad H. Danesh, Maxime Wabartha, Stanley Wu +2

Deploying reinforcement learning (RL) policies in real-world involves significant challenges, including distribution shifts, safety concerns, and the impracticality of direct inter…

cs.LG2025

YRC-Bench: A Benchmark for Learning to Coordinate with Experts

Mohamad H. Danesh, Nguyen X. Khanh, Tu Trinh +1

When deployed in the real world, AI agents will inevitably face challenges that exceed their individual capabilities. A critical component of AI safety is an agent's ability to rec…

cs.LG2024

Getting By Goal Misgeneralization With a Little Help From a Mentor

Tu Trinh, Mohamad H. Danesh, Nguyen X. Khanh +1

While reinforcement learning (RL) agents often perform well during training, they can struggle with distribution shift in real-world deployments. One particularly severe risk of di…

cs.LG2024★ 3 cited

Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning

Shengyi Huang, Quentin Gallouédec, Florian Felten +30

In many Reinforcement Learning (RL) papers, learning curves are useful indicators to measure the effectiveness of RL algorithms. However, the complete raw data of the learning curv…