3 papers
cs.AI2025
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
Syrine Belakaria, Joshua Kazdan, Charles Marx +5
Reinforcement learning from human feedback (RLHF) has become a cornerstone of the training and alignment pipeline for large language models (LLMs). Recent advances, such as direct…
cs.LG2025
Preference-Guided Diffusion for Multi-Objective Offline Optimization
Yashas Annadani, Syrine Belakaria, Stefano Ermon +2
Offline multi-objective optimization aims to identify Pareto-optimal solutions given a dataset of designs and their objective values. In this work, we propose a preference-guided d…
cs.LG2024
Non-Myopic Multi-Objective Bayesian Optimization
Syrine Belakaria, Alaleh Ahmadianshalchi, Barbara Engelhardt +2
We consider the problem of finite-horizon sequential experimental design to solve multi-objective optimization (MOO) of expensive black-box objective functions. This problem arises…