6 papers
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
Yuchen Zhu, Yufeng Zhang, Zhaoran Wang +2
This paper studies minimax optimization problems defined over infinite-dimensional function classes of overparameterized two-layer neural networks. In particular, we consider the m…
An Analysis of Attention via the Lens of Exchangeability and Latent Variable Models
Yufeng Zhang, Boyi Liu, Qi Cai +2
With the attention mechanism, transformers achieve significant empirical successes. Despite the intuitive understanding that transformers perform relational inference over long seq…
Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic
Yufeng Zhang, Siyu Chen, Zhuoran Yang +2
Actor-critic (AC) algorithms, empowered by neural networks, have had significant empirical success in recent years. However, most of the existing theoretical support for AC algorit…
Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
Zhuoran Yang, Yufeng Zhang, Yongxin Chen +1
We consider the optimization problem of minimizing a functional defined over a family of probability distributions, where the objective functional is assumed to possess a variation…
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
Yufeng Zhang, Qi Cai, Zhuoran Yang +2
Temporal-difference and Q-learning play a key role in deep reinforcement learning, where they are empowered by expressive nonlinear function approximators such as neural networks.…
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
Hongyi Guo, Zhihan Liu, Yufeng Zhang +1
Large Language Models (LLMs) harness extensive data from the Internet, storing a broad spectrum of prior knowledge. While LLMs have proven beneficial as decision-making aids, their…