8 papers
GroupDPO: Memory efficient Group-wise Direct Preference Optimization
Jixuan Leng, Si Si, Hsiang-Fu Yu +2
Preference optimization is widely used to align Large Language Models (LLMs) with preference feedback. However, most existing methods train on a single positive-negative pair per p…
Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation
Seamus Somerstep, Vinod Raman, Unique Subedi +1
Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. The first, referred to as su…
Missing Mass for Differentially Private Domain Discovery
Travis Dick, Matthew Joseph, Vinod Raman
We study several problems in differentially private domain discovery, where each user holds a subset of items from a shared but unknown domain, and the goal is to output an informa…
On Generation in Metric Spaces
Jiaxun Li, Vinod Raman, Ambuj Tewari
We study generation in separable metric instance spaces. We extend the language generation framework from Kleinberg and Mullainathan [2024] beyond countable domains by defining nov…
The Complexity of Sequential Prediction in Dynamical Systems
Vinod Raman, Unique Subedi, Ambuj Tewari
We study the problem of learning to predict the next state of a dynamical system when the underlying evolution function is unknown. Unlike previous work, we place no parametric ass…
Generation through the lens of learning theory
Jiaxun Li, Vinod Raman, Ambuj Tewari
We study generation through the lens of statistical learning theory. First, we abstract and formalize the results of Gold [1967], Angluin [1979], Angluin [1980] and Kleinberg and M…