2 papers
cs.LG2024
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
Zihao Li, Yuan Cao, Cheng Gao +5
Transformers have achieved great success in recent years. Interestingly, transformers have shown particularly strong in-context learning capability -- even without fine-tuning, the…
stat.ML2024
Global Convergence in Training Large-Scale Transformers
Cheng Gao, Yuan Cao, Zihao Li +5
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously an…