15 citations · 25 across the 20 of their papers we have counts for
7 papers · 1 filter
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
Elham Daneshmand, Majid Khadiv, Glen Berseth +1
Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the go…
Drift Q-Learning
Anas Houssaini, Mohamad H. Danesh, Amin Abyaneh +3
Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies h…
Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations
Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh +4
Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stoc…
SAD-Flower: Flow Matching for Safe, Admissible, and Dynamically Consistent Planning
Tzu-Yuan Huang, Armin Lederer, Dai-Jie Wu +6
Flow matching (FM) has shown promising results in data-driven planning. However, it inherently lacks formal guarantees for ensuring state and action constraints, whose satisfaction…
Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation
Mohamad H. Danesh, Maxime Wabartha, Stanley Wu +2
Deploying reinforcement learning (RL) policies in real-world involves significant challenges, including distribution shifts, safety concerns, and the impracticality of direct inter…
Contractive Dynamical Imitation Policies for Efficient Out-of-Sample Recovery
Amin Abyaneh, Mahrokh G. Boroujeni, Hsiu-Chin Lin +1
Imitation learning is a data-driven approach to learning policies from expert behavior, but it is prone to unreliable outcomes in out-of-sample (OOS) regions. While previous resear…