2 papers
cs.AI2026
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
cs.LG2026
MixDPO: Modeling Preference Strength for Pluralistic Alignment
Saki Imai, Pedram Heydari, Anthony Sicilia +3
Preference based alignment objectives implicitly assume that all human preferences are expressed with equal strength. In practice, however, preference strength varies across indivi…