activity
20232025
most citedLocoMuJoCo: A Comprehensive Imitation Learning Benchmark for Locomotion

2 citations · 3 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2025

Discrete Variational Autoencoding via Policy Search

Michael Drolet, Firas Al-Hafez, Aditya Bhatt +2

Discrete latent bottlenecks in variational autoencoders (VAEs) offer high bit efficiency and can be modeled with autoregressive discrete distributions, enabling parameter-efficient…

cs.RO2024

The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control

Oleg Kaidanov, Firas Al-Hafez, Yusuf Suvari +2

Humanoids have the potential to be the ideal embodiment in environments designed for humans. Thanks to the structural similarity to the human body, they benefit from rich sources o…

cs.RO2024

Exciting Action: Investigating Efficient Exploration for Learning Musculoskeletal Humanoid Locomotion

Henri-Jacques Geiß, Firas Al-Hafez, Andre Seyfarth +2

Learning a locomotion controller for a musculoskeletal system is challenging due to over-actuation and high-dimensional action space. While many reinforcement learning methods atte…

cs.LG2023★ 2 cited

LocoMuJoCo: A Comprehensive Imitation Learning Benchmark for Locomotion

Firas Al-Hafez, Guoping Zhao, Jan Peters +1

Imitation Learning (IL) holds great promise for enabling agile locomotion in embodied agents. However, many existing locomotion benchmarks primarily focus on simplified toy tasks,…

cs.LG2023

Time-Efficient Reinforcement Learning with Stochastic Stateful Policies

Firas Al-Hafez, Guoping Zhao, Jan Peters +1

Stateful policies play an important role in reinforcement learning, such as handling partially observable environments, enhancing robustness, or imposing an inductive bias directly…

cs.LG2023★ 1 cited

LS-IQ: Implicit Reward Regularization for Inverse Reinforcement Learning

Firas Al-Hafez, Davide Tateo, Oleg Arenz +2

Recent methods for imitation learning directly learn a -function using an implicit reward formulation rather than an explicit reward function. However, these methods generally r…