activity
20162026
most citedSingle-shot Foothold Selection and Constraint Evaluation for Quadruped Locomotion

15 citations · 25 across the 20 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

Elham Daneshmand, Majid Khadiv, Glen Berseth +1

Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the go…

cs.LG2026

Drift Q-Learning

Anas Houssaini, Mohamad H. Danesh, Amin Abyaneh +3

Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies h…

cs.LG2026

Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations

Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh +4

Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stoc…

cs.LG2025

SAD-Flower: Flow Matching for Safe, Admissible, and Dynamically Consistent Planning

Tzu-Yuan Huang, Armin Lederer, Dai-Jie Wu +6

Flow matching (FM) has shown promising results in data-driven planning. However, it inherently lacks formal guarantees for ensuring state and action constraints, whose satisfaction…

cs.LG2025

Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation

Mohamad H. Danesh, Maxime Wabartha, Stanley Wu +2

Deploying reinforcement learning (RL) policies in real-world involves significant challenges, including distribution shifts, safety concerns, and the impracticality of direct inter…

cs.LG2024

Contractive Dynamical Imitation Policies for Efficient Out-of-Sample Recovery

Amin Abyaneh, Mahrokh G. Boroujeni, Hsiu-Chin Lin +1

Imitation learning is a data-driven approach to learning policies from expert behavior, but it is prone to unreliable outcomes in out-of-sample (OOS) regions. While previous resear…