4 papers
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Matthew Nguyen, Kyle Cox, Austin Meek +1
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is p…
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workf…
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other mode…
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
Matthew Bozoukov, Matthew Nguyen, Shubkarman Singh +2
Recent studies have revealed that LLMs can exhibit behavioral self-awareness: the ability to accurately describe or predict their own learned behaviors without explicit supervision…