3 papers
cs.CL2026
Rushes: A Human Preference Dataset for Pluralistic Alignment
Michael Xu, Jorge Leandro, Sudha Rao +5
We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface…
cs.IR2026
Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers
Yue Kang, Zhuoyi Huang, Benji Schussheim +19
In enterprise search, building high-quality datasets at scale remains a central challenge due to the difficulty of acquiring labeled data. To resolve this challenge, we propose an…
cs.CL2025
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
Yao Dou, Michel Galley, Baolin Peng +6
Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn convers…