5.7k citations · 5.7k across the 2 of their papers we have counts for
4 papers
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
Zihao Li, Yuan Cao, Cheng Gao +5
Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, the…
Global Convergence in Training Large-Scale Transformers
Cheng Gao, Yuan Cao, Zihao Li +5
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously an…
Learn to Cluster Faces with Better Subgraphs
Yuan Cao, Di Jiang, Guanqun Hou +3
Face clustering can provide pseudo-labels to the massive unlabeled face data and improve the performance of different face recognition models. The existing clustering methods gener…
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen +28
Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based tr…