12 papers
Optimistic Online LQR via Intrinsic Rewards
Marcell Bartos, Bruce D. Lee, Lenart Treven +4
Optimism in the face of uncertainty is a popular approach to balance exploration and exploitation in reinforcement learning. Here, we consider the online linear quadratic regulator…
Sample-efficient and Scalable Exploration in Continuous-Time RL
Klemens Iten, Lenart Treven, Bhavya Sukhija +2
Reinforcement learning algorithms are typically designed for discrete-time dynamics, even though the underlying real-world control systems are often continuous in time. In this pap…
SOMBRL: Scalable and Optimistic Model-Based RL
Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza +3
We address the challenge of efficient exploration in model-based reinforcement learning (MBRL), where the system dynamics are unknown and the RL agent must learn directly from onli…
TARC: Time-Adaptive Robotic Control
Arnav Sukhija, Lenart Treven, Jin Cheng +3
Fixed-frequency control in robotics imposes a trade-off between the efficiency of low-frequency control and the robustness of high-frequency control, a limitation not seen in adapt…
SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer
Yarden As, Chengrui Qu, Benjamin Unger +6
Deploying reinforcement learning (RL) safely in the real world is challenging, as policies trained in simulators must face the inevitable sim-to-real gap. Robust safe RL techniques…
Simulation Priors for Data-Efficient Deep Learning
Lenart Treven, Bhavya Sukhija, Jonas Rothfuss +3
How do we enable AI systems to efficiently learn in the real-world? First-principles models are widely used to simulate natural systems, but often fail to capture real-world comple…