1 citations · 2 across the 3 of their papers we have counts for
3 papers
On the Power of Convolution Augmented Transformer
Mingchen Li, Xuechen Zhang, Yixiao Huang +1
The transformer architecture has catalyzed revolutionary advances in language modeling. However, recent architectural recipes, such as state-space models, have bridged the performa…
Mechanics of Next Token Prediction with Self-Attention
Yingcong Li, Yixiao Huang, M. Emrullah Ildiz +2
Transformer-based language models are trained on large datasets to predict the next token given an input sequence. Despite this simple training objective, they have led to revoluti…
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
M. Emrullah Ildiz, Yixiao Huang, Yingcong Li +2
Modern language models rely on the transformer architecture and attention mechanism to perform language understanding and text generation. In this work, we study learning a 1-layer…