2 papers
cs.LG2026
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
Elham Daneshmand, Majid Khadiv, Glen Berseth +1
Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the go…
cs.RO2026
SLowRL: Safe Low-Rank Adaptation Reinforcement Learning for Locomotion
Elham Daneshmand, Shafeef Omar, Glen Berseth +2
Sim-to-real transfer of locomotion policies often leads to performance degradation due to the inevitable sim-to-real gap. Naively fine-tuning these policies directly on hardware is…