4 papers
Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems
Yaochen Zhu, Harald Steck, James McInerney +4
Direct preference optimization (DPO) is a simple and effective alignment strategy for large language models (LLMs) based on pairwise preferences. In recommender systems, however, u…
Entropy After </Think> for reasoning model early exiting
Xi Wang, James McInerney, Lequn Wang +1
Reasoning LLMs show improved performance with longer chains of thought. However, recent work has highlighted their tendency to overthink, continuing to revise answers even after re…
Optimization of Epsilon-Greedy Exploration
Ethan Che, Hakan Ceylan, James McInerney +1
Modern recommendation systems rely on exploration to learn user preferences for new items, typically implementing uniform exploration policies (e.g., epsilon-greedy) due to their s…
Variation Due to Regularization Tractably Recovers Bayesian Deep Learning
James McInerney, Nathan Kallus
Uncertainty quantification in deep learning is crucial for safe and reliable decision-making in downstream tasks. Existing methods quantify uncertainty at the last layer or other a…