5 papers
Global Convergence and Rich Feature Learning in -Layer Infinite-Width Neural Networks under P Parametrization
Zixiang Chen, Greg Yang, Qingyue Zhao +1
Despite deep neural networks' powerful representation learning capabilities, theoretical understanding of how networks can simultaneously achieve meaningful feature learning and gl…
Convergence of Score-Based Discrete Diffusion Models: A Discrete-Time Analysis
Zikun Zhang, Zixiang Chen, Quanquan Gu
Diffusion models have achieved great success in generating high-dimensional samples across various applications. While the theoretical guarantees for continuous-state diffusion mod…
Matching the Statistical Query Lower Bound for -Sparse Parity Problems with Sign Stochastic Gradient Descent
Yiwen Kou, Zixiang Chen, Quanquan Gu +1
The -sparse parity problem is a classical problem in computational complexity and algorithmic theory, serving as a key benchmark for understanding computational classes. In this…
Fast Sampling via Discrete Non-Markov Diffusion Models with Predetermined Transition Time
Zixiang Chen, Huizhuo Yuan, Yongqian Li +3
Discrete diffusion models have emerged as powerful tools for high-quality data generation. Despite their success in discrete spaces, such as text generation tasks, the acceleration…
Self-Play Preference Optimization for Language Model Alignment
Yue Wu, Zhiqing Sun, Huizhuo Yuan +3
Standard reinforcement learning from human feedback (RLHF) approaches relying on parametric models like the Bradley-Terry model fall short in capturing the intransitivity and irrat…