4 papers
Abstract Markov Random Fields
Leon Lang, Clélia de Mulatier, Rick Quax +1
Markov random fields are known to be fully characterized by properties of their information diagrams, or I-diagrams. In particular, for Markov random fields, regions in the I-diagr…
Modeling Human Beliefs about AI Behavior for Scalable Oversight
Leon Lang, Patrick Forré
As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators m…
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
Lukas Fluri, Leon Lang, Alessandro Abate +3
In reinforcement learning, specifying reward functions that capture the intended task can be very challenging. Reward learning aims to address this issue by learning the reward fun…
Factored space models: Towards causality between levels of abstraction
Scott Garrabrant, Matthias Georg Mayer, Magdalena Wache +3
Causality plays an important role in understanding intelligent behavior, and there is a wealth of literature on mathematical models for causality, most of which is focused on causa…