253 citations · 628 across the 9 of their papers we have counts for
9 papers
Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
Yiping Lu, Zhuohan Li, Di He +5
The Transformer architecture is widely used in natural language processing. Despite its success, the design principle of the Transformer remains elusive. In this paper, we provide…
Distributed Bandit Learning: Near-Optimal Regret with Efficient Communication
Yuanhao Wang, Jiachen Hu, Xiaoyu Chen +1
We study the problem of regret minimization for distributed bandits learning, in which agents work collaboratively to minimize their total regret under the coordination of a ce…
The Expressive Power of Neural Networks: A View from the Width
Zhou Lu, Hongming Pu, Feicheng Wang +2
The expressive power of neural networks is important for understanding deep learning. Most existing works consider this problem from the view of the depth of a network. In this pap…
Generalization Bounds of SGLD for Non-convex Learning: Two Theoretical Viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai +1
Algorithm-dependent generalization error bounds are central to statistical learning theory. A learning algorithm may use a large hypothesis space, but the limited number of iterati…
Zero-Shot Fine-Grained Classification by Deep Feature Learning with Semantics
Aoxue Li, Zhiwu Lu, Liwei Wang +3
Fine-grained image classification, which aims to distinguish images with subtle distinctions, is a challenging task due to two main issues: lack of sufficient training data for eve…
Collect at Once, Use Effectively: Making Non-interactive Locally Private Learning Possible
Kai Zheng, Wenlong Mou, Liwei Wang
Non-interactive Local Differential Privacy (LDP) requires data analysts to collect data from users through noisy channel at once. In this paper, we extend the frontiers of Non-inte…