Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Modeling Human Beliefs about AI Behavior for Scalable Oversight
Leon Lang, Patrick Forré
As AI systems advance beyond human capabilities, scalable oversight becomes critical: how can we supervise AI that exceeds our abilities? A key challenge is that human evaluators m…
cs.AI2024
Factored space models: Towards causality between levels of abstraction
Scott Garrabrant, Matthias Georg Mayer, Magdalena Wache +3
Causality plays an important role in understanding intelligent behavior, and there is a wealth of literature on mathematical models for causality, most of which is focused on causa…