3 papers
cs.LG2026
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
Amirhossein Afsharrad, Ruida Zhou, Luca Viano +2
Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal prefere…
cs.LG2025
Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
Emi Soroka, Tanmay Chopra, Krish Desai +1
Large language models (LLMs) have seen increasing popularity in enterprise applications where AI agents and humans engage in objective-driven interactions. However, these systems a…
cs.LG2025
Learning Temporal Logic Predicates from Data with Statistical Guarantees
Emi Soroka, Rohan Sinha, Sanjay Lall
Temporal logic rules are often used in control and robotics to provide structured, human-interpretable descriptions of trajectory data. These rules have numerous applications inclu…