7 papers
AdaBoN: Adaptive Best-of-N Alignment
Vinod Raman, Hilal Asi, Satyen Kale
Recent advances in test-time alignment methods, such as Best-of-N sampling, offer a simple and effective way to steer language models (LMs) toward preferred behaviors using reward…
Transductive and Learning-Augmented Online Regression
Vinod Raman, Shenghao Xie, Samson Zhou
Motivated by the predictable nature of real-life in data streams, we study online regression when the learner has access to predictions about future examples. In the extreme case,…
Optimal Stopping vs Best-of- for Inference Time Optimization
Yusuf Kalayci, Vinod Raman, Shaddin Dughmi
Large language model (LLM) generation often requires balancing output quality against inference cost, especially when using multiple generations. We introduce a new framework for i…
Generation from Noisy Examples
Ananth Raman, Vinod Raman
We continue to study the learning-theoretic foundations of generation by extending the results from Kleinberg and Mullainathan [2024] and Li et al. [2024] to account for noisy exam…
Representative Language Generation
Charlotte Peale, Vinod Raman, Omer Reingold
We introduce "representative generation," extending the theoretical framework for generation proposed by Kleinberg et al. (2024) and formalized by Li et al. (2024), to additionally…
Faster Rates for Private Adversarial Bandits
Hilal Asi, Vinod Raman, Kunal Talwar
We design new differentially private algorithms for the problems of adversarial bandits and bandits with expert advice. For adversarial bandits, we give a simple and efficient conv…