3 citations · 5 across the 2 of their papers we have counts for
4 papers
Persona Features Control Emergent Misalignment
Miles Wang, Tom Dupré la Tour, Olivia Watkins +8
Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that…
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren +3
The ability to communicate uncertainty, risk, and limitation is crucial for the safety of large language models. However, current evaluations of these abilities rely on simple cali…
modeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models
Nathan A. Chi, Teodor Malchev, Riley Kong +5
We introduce modeLing, a novel benchmark of Linguistics Olympiad-style puzzles which tests few-shot reasoning in AI systems. Solving these puzzles necessitates inferring aspects of…
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
Andrew M. Bean, Simi Hellsten, Harry Mayne +5
In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we eval…