2 papers
cs.LG2026
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
Amirhossein Afsharrad, Ruida Zhou, Luca Viano +2
Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal prefere…
cs.LG2025
Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
Emi Soroka, Tanmay Chopra, Krish Desai +1
Large language models (LLMs) have seen increasing popularity in enterprise applications where AI agents and humans engage in objective-driven interactions. However, these systems a…