activity
20182026
most citedSpatial Computing and Intuitive Interaction: Bringing Mixed Reality and Robotics Together

68 citations · 181 across the 45 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2026

Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning

Jiajun Hu, Nuria Armengol Urpi, Jin Cheng +1

Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies for any reward function directly at test…

cs.LG2026

Safe Exploration via Policy Priors

Manuel Wendl, Yarden As, Manish Prajapat +3

Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated) environments. In this work, we tackle thi…

cs.LG2025

Simulation Priors for Data-Efficient Deep Learning

Lenart Treven, Bhavya Sukhija, Jonas Rothfuss +3

How do we enable AI systems to efficiently learn in the real-world? First-principles models are widely used to simulate natural systems, but often fail to capture real-world comple…

cs.LG2024★ 1 cited

MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization

Bhavya Sukhija, Stelian Coros, Andreas Krause +2

Reinforcement learning (RL) algorithms aim to balance exploiting the current best strategy with exploring new options that could lead to higher rewards. Most common RL algorithms u…

cs.LG2024

ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning

Yarden As, Bhavya Sukhija, Lenart Treven +3

Reinforcement learning (RL) is ubiquitous in the development of modern AI systems. However, state-of-the-art RL agents require extensive, and potentially unsafe, interactions with…

cs.LG2024

NeoRL: Efficient Exploration for Nonepisodic RL

Bhavya Sukhija, Lenart Treven, Florian Dörfler +2

We study the problem of nonepisodic reinforcement learning (RL) for nonlinear dynamical systems, where the system dynamics are unknown and the RL agent has to learn from a single t…