2 papers
cs.LG2025
The Limits of Preference Data for Post-Training
Eric Zhao, Jessica Dai, Pranjal Awasthi
Recent progress in strengthening the capabilities of large language models has stemmed from applying reinforcement learning to domains with automatically verifiable outcomes. A key…
cs.LG2024
Learning With Multi-Group Guarantees For Clusterable Subpopulations
Jessica Dai, Nika Haghtalab, Eric Zhao
A canonical desideratum for prediction problems is that performance guarantees should hold not just on average over the population, but also for meaningful subpopulations within th…