7 papers
Drift Q-Learning
Anas Houssaini, Mohamad H. Danesh, Amin Abyaneh +3
Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies h…
Morphology-Conditioned World Model for Cross-Embodiment Quadrupedal Locomotion
Mohamad H. Danesh, Chenhao Li, Amin Abyaneh +5
World models promise a paradigm shift in robotics, where an agent learns the physics of its environment once and then acquires behaviors efficiently. Yet the learned dynamics model…
Contractive Diffusion Policies: Robust Action Diffusion via Contractive Score-Based Sampling with Differential Equations
Amin Abyaneh, Charlotte Morissette, Mohamad H. Danesh +4
Diffusion policies have emerged as powerful generative models for offline policy learning, whose sampling process can be rigorously characterized by a score function guiding a stoc…
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
Mohamad H. Danesh, Nguyen X. Khanh, Tu Trinh +1
When deployed in the real world, AI agents will inevitably face challenges that exceed their individual capabilities. A critical component of AI safety is an agent's ability to rec…
VOCALoco: Viability-Optimized Cost-aware Adaptive Locomotion
Stanley Wu, Mohamad H. Danesh, Simon Li +5
Recent advancements in legged robot locomotion have facilitated traversal over increasingly complex terrains. Despite this progress, many existing approaches rely on end-to-end dee…
Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation
Mohamad H. Danesh, Maxime Wabartha, Stanley Wu +2
Deploying reinforcement learning (RL) policies in real-world involves significant challenges, including distribution shifts, safety concerns, and the impracticality of direct inter…