3 papers
cs.LG2026
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
Wei Chen, Yubing Wu, Junmei Yang +5
Preference optimization is widely used to align large language models (LLMs) with human preferences. However, many margin-based methods also suppress the chosen response when they…
stat.ML2024
Efficient Online Set-valued Classification with Bandit Feedback
Zhou Wang, Xingye Qiao
Conformal prediction is a distribution-free method that wraps a given machine learning model and returns a set of plausible labels that contain the true label with a prescribed cov…
cs.IR2023
AutoML for Large Capacity Modeling of Meta's Ranking Systems
Hang Yin, Kuang-Hung Liu, Mengying Sun +16
Web-scale ranking systems at Meta serving billions of users is complex. Improving ranking models is essential but engineering heavy. Automated Machine Learning (AutoML) can release…