13 papers
MORFEO control strategy
Guido Agapito, Lorenzo Busoni, Cédric Plantet +7
The ESO Extremely Large Telescope (ELT) will offer unprecedented sensitivity and resolution in the near-infrared, marking a new era for ground-based astronomy. Among its key imagin…
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
Max Lamparth, Daniel Fein, Andreas Haupt +2
Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than…
Behavior-Consistent Deep Reinforcement Learning
Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker +2
Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. I…
Iterative Compositional Data Generation for Robot Control
Anh-Quan Pham, Marcel Hussing, Shubhankar P. Patankar +3
Collecting robotic manipulation data is expensive, making it impractical to acquire demonstrations for the combinatorially large space of tasks that arise in multi-object, multi-ro…
Replicable Reinforcement Learning with Linear Function Approximation
Eric Eaton, Marcel Hussing, Michael Kearns +3
Replication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learnin…
Relative Entropy Pathwise Policy Optimization
Claas Voelcker, Axel Brunnbauer, Marcel Hussing +6
Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their high variance often undermines tr…