5 papers
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workf…
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other mode…
Measuring Weak-to-Strong Legibility of Reasoning Models
Dani Roytburg, Shreya Sridhar, Daphne Ippolito
Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such as inter-model monitoring or distilla…
Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research
Dani Roytburg, Beck Miller
While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "alig…
Words and Action: Modeling Linguistic Leadership in #BlackLivesMatter Communities
Dani Roytburg, Deborah Olorunisola, Sandeep Soni +1
In this project, we describe a method of modeling semantic leadership across a set of communities associated with the #BlackLivesMatter movement, which has been informed by qualita…