3 papers
cs.AI2026
Which Pairs to Compare for LLM Post-Training?
Jiangze Han, Vineet Goyal, Will Ma
Preference-based post-training has become a central paradigm for aligning language models. A common data-collection strategy is to generate a small set of completions for each prom…
cs.LG2025
Collaborative Min-Max Regret in Grouped Multi-Armed Bandits
Moïse Blanchard, Vineet Goyal
We study the impact of sharing exploration in multi-armed bandits in a grouped setting where a set of groups have overlapping feasible action sets [Baek and Farias '24]. In this gr…
math.OC2024
Distributionally Robust Newsvendor on a Metric
Ayoub Foussoul, Vineet Goyal
We consider a fundamental generalization of the classical newsvendor problem where the seller needs to decide on the inventory of a product jointly for multiple locations on a metr…