1.6k citations · 1.9k across the 32 of their papers we have counts for
33 papers
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
Jiawei Lin, Yuanlong Li, Guokai Chen +1
Transformer models rely heavily on the scaled dot-product attention (SDPA) operation, typically implemented as FlashAttention. Characterized by its frequent interleaving of matrix…
Textbooks Are All You Need II: phi-1.5 technical report
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan +3
We continue the investigation into the power of smaller Transformer-based language models as initiated by \textbf{TinyStories} -- a 10 million parameter model that can produce cohe…
Efficient RLHF: Reducing the Memory Usage of PPO
Michael Santacroce, Yadong Lu, Han Yu +2
Reinforcement Learning with Human Feedback (RLHF) has revolutionized language modeling by aligning models with human preferences. However, the RL stage, Proximal Policy Optimizatio…
Generating Faithful Text From a Knowledge Graph with Noisy Reference Text
Tahsina Hashem, Weiqing Wang, Derry Tanti Wijaya +2
Knowledge Graph (KG)-to-Text generation aims at generating fluent natural-language text that accurately represents the information of a given knowledge graph. While significant pro…
The Implicit Bias of Batch Normalization in Linear Models and Two-layer Linear Convolutional Neural Networks
Yuan Cao, Difan Zou, Yuanzhi Li +1
We study the implicit bias of batch normalization trained by gradient descent. We show that when learning a linear model with batch normalization for binary classification, gradien…
Length Generalization in Arithmetic Transformers
Samy Jelassi, Stéphane d'Ascoli, Carles Domingo-Enrich +3
We examine how transformers cope with two challenges: learning basic integer arithmetic, and generalizing to longer sequences than seen during training. We find that relative posit…