42 citations · 138 across the 25 of their papers we have counts for
21 papers · 1 filter
Conversational Dueling Bandits in Generalized Linear Models
Shuhua Yang, Hui Yuan, Xiaoying Zhang +3
Conversational recommendation systems elicit user preferences by interacting with users to obtain their feedback on recommended commodities. Such systems utilize a multi-armed band…
Contractual Reinforcement Learning: Pulling Arms with Invisible Hands
Jibang Wu, Siyu Chen, Mengdi Wang +2
The agency problem emerges in today's large scale machine learning tasks, where the learners are unable to direct content creation or enforce data collection. In this work, we prop…
Provable Statistical Rates for Consistency Diffusion Models
Zehao Dou, Minshuo Chen, Mengdi Wang +1
Diffusion models have revolutionized various application domains, including computer vision and audio generation. Despite the state-of-the-art performance, diffusion models are kno…
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
Mucong Ding, Souradip Chakraborty, Vibhu Agrawal +5
Reinforcement Learning from Human Feedback (RLHF) is a key method for aligning large language models (LLMs) with human preferences. However, current offline alignment approaches li…
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
Xiang Ji, Sanjeev Kulkarni, Mengdi Wang +1
This work studies the challenge of aligning large language models (LLMs) with offline preference data. We focus on alignment by Reinforcement Learning from Human Feedback (RLHF) in…
An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization
Minshuo Chen, Song Mei, Jianqing Fan +1
Diffusion models, a powerful and universal generative AI technology, have achieved tremendous success in computer vision, audio, reinforcement learning, and computational biology.…