3 papers
stat.ML2026
Reward Learning from Best-of- Preference Data: Targets, Tradeoffs, and Design Principles
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Best-of- sampling is widely used to construct pairwise preference data: candidates are drawn from a base distribution, and the best is paired with a rejected response. Despi…
cs.LG2026
What Does Preference Learning Recover from Pairwise Comparison Data?
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical dataset consists of triplets $(x…
cs.LG2025
Learning from Interval Targets
Rattana Pukdee, Ziqi Ke, Chirag Gupta
We study the problem of regression with interval targets, where only upper and lower bounds on target values are available in the form of intervals. This problem arises when the ex…