collaborators

5 papers

cs.CL2026

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workf…

cs.CL2026

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other mode…

cs.MA2026

Measuring Weak-to-Strong Legibility of Reasoning Models

Dani Roytburg, Shreya Sridhar, Daphne Ippolito

Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such as inter-model monitoring or distilla…

cs.AI2025

Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research

Dani Roytburg, Beck Miller

While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "alig…

cs.CL2025

Words and Action: Modeling Linguistic Leadership in #BlackLivesMatter Communities

Dani Roytburg, Deborah Olorunisola, Sandeep Soni +1

In this project, we describe a method of modeling semantic leadership across a set of communities associated with the #BlackLivesMatter movement, which has been informed by qualita…