6 papers · 1 filter
Oracle-Robust Online Alignment for Large Language Models
Zimeng Li, Mudit Gaur, Vaneet Aggarwal
We study online alignment of large language models under misspecified preference feedback, where the observed preference oracle deviates from an ideal but unknown ground-truth orac…
Order-Optimal Sample Complexity of Rectified Flows
Hari Krishna Sahoo, Mudit Gaur, Vaneet Aggarwal
Recently, flow-based generative models have shown superior efficiency compared to diffusion models. In this paper, we study rectified flow models, which constrain transport traject…
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
Mudit Gaur, Prashant Trivedi, Shuchin Aeron +3
Flow matching has recently emerged as a promising alternative to diffusion-based generative models, offering faster sampling and simpler training by learning continuous flows gover…
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
Mudit Gaur, Prashant Trivedi, Sasidhar Kunapuli +2
Diffusion models have demonstrated state-of-the-art performance across vision, language, and scientific domains. Despite their empirical success, prior theoretical analyses of the…
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
Mudit Gaur, Utsav Singh, Amrit Singh Bedi +2
Bilevel reinforcement learning (BRL) has emerged as a powerful framework for aligning generative models, yet its theoretical foundations, especially sample complexity bounds, remai…
On The Global Convergence Of Online RLHF With Neural Parametrization
Mudit Gaur, Amrit Singh Bedi, Raghu Pasupathy +1
The importance of Reinforcement Learning from Human Feedback (RLHF) in aligning large language models (LLMs) with human values cannot be overstated. RLHF is a three-stage process t…