7 citations · 7 across the 3 of their papers we have counts for
3 papers
Soft Preference Optimization: Aligning Language Models to Expert Distributions
Arsalan Sharifnassab, Saber Salehkaleybar, Sina Ghiassian +2
We propose Soft Preference Optimization (SPO), a method for aligning generative models, such as Large Language Models (LLMs), with human preferences, without the need for a reward…
Automatic Music Playlist Generation via Simulation-based Reinforcement Learning
Federico Tomasi, Joseph Cauteruccio, Surya Kanoria +3
Personalization of playlists is a common feature in music streaming services, but conventional techniques, such as collaborative filtering, rely on explicit assumptions regarding c…
What to Learn, and How: Toward Effective Learning from Rationales
Samuel Carton, Surya Kanoria, Chenhao Tan
Learning from rationales seeks to augment model prediction accuracy using human-annotated rationales (i.e. subsets of input tokens) that justify their chosen labels, often in the f…