2 papers
cs.AI2026
Demonstrating Generalization Failures via Mixtures of Conditional Policies
Jou Barzdukas, Jack Peck, Julian Schulz +3
Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes…
cs.CL2026
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other mode…