4 papers
Which Pairs to Compare for LLM Post-Training?
Jiangze Han, Vineet Goyal, Will Ma
Preference-based post-training has become a central paradigm for aligning language models. A common data-collection strategy is to generate a small set of completions for each prom…
PREFER: Personalized Review Summarization with Online Preference Learning
Millend Roy, Agostino Capponi, Vineet Goyal
Product reviews significantly influence purchasing decisions on e-commerce platforms. However, the sheer volume of reviews can overwhelm users, obscuring the information most relev…
Collaborative Min-Max Regret in Grouped Multi-Armed Bandits
Moïse Blanchard, Vineet Goyal
We study the impact of sharing exploration in multi-armed bandits in a grouped setting where a set of groups have overlapping feasible action sets [Baek and Farias '24]. In this gr…
Distributionally Robust Newsvendor on a Metric
Ayoub Foussoul, Vineet Goyal
We consider a fundamental generalization of the classical newsvendor problem where the seller needs to decide on the inventory of a product jointly for multiple locations on a metr…