2 citations · 5 across the 14 of their papers we have counts for
4 papers · 1 filter
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
Botos Csaba, Sreejan Kumar, Austin Tudor David Andrews +6
Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems lea…
When Do LLM Preferences Predict Downstream Behavior?
Katarina Slama, Alexandra Souly, Dishank Bansal +3
Preference-driven behavior in LLMs may be a necessary precondition for AI misalignment such as sandbagging: models cannot strategically pursue misaligned goals unless their behavio…
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
Lennart Luettgau, Harry Coppock, Magda Dubois +2
As Large Language Models (LLMs) and other AI systems evolve, robustly estimating their capabilities from inherently stochastic outputs while systematically quantifying uncertainty…
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Christopher Summerfield, Lennart Luettgau, Magda Dubois +9
We examine recent research that asks whether current AI systems may be developing a capacity for "scheming" (covertly and strategically pursuing misaligned goals). We compare curre…