activity
20202025
most citedSparks of Artificial General Intelligence: Early experiments with GPT-4

1.6k citations · 1.9k across the 32 of their papers we have counts for

collaborators

33 papers

cs.AR2025

SystolicAttention: Fusing FlashAttention within a Single Systolic Array

Jiawei Lin, Yuanlong Li, Guokai Chen +1

Transformer models rely heavily on the scaled dot-product attention (SDPA) operation, typically implemented as FlashAttention. Characterized by its frequent interleaving of matrix…

cs.CL2023★ 52 cited

Textbooks Are All You Need II: phi-1.5 technical report

Yuanzhi Li, Sébastien Bubeck, Ronen Eldan +3

We continue the investigation into the power of smaller Transformer-based language models as initiated by \textbf{TinyStories} -- a 10 million parameter model that can produce cohe…

cs.LG2023★ 2 cited

Efficient RLHF: Reducing the Memory Usage of PPO

Michael Santacroce, Yadong Lu, Han Yu +2

Reinforcement Learning with Human Feedback (RLHF) has revolutionized language modeling by aligning models with human preferences. However, the RL stage, Proximal Policy Optimizatio…

cs.CL2023

Generating Faithful Text From a Knowledge Graph with Noisy Reference Text

Tahsina Hashem, Weiqing Wang, Derry Tanti Wijaya +2

Knowledge Graph (KG)-to-Text generation aims at generating fluent natural-language text that accurately represents the information of a given knowledge graph. While significant pro…

cs.LG2023

The Implicit Bias of Batch Normalization in Linear Models and Two-layer Linear Convolutional Neural Networks

Yuan Cao, Difan Zou, Yuanzhi Li +1

We study the implicit bias of batch normalization trained by gradient descent. We show that when learning a linear model with batch normalization for binary classification, gradien…

cs.LG2023★ 2 cited

Length Generalization in Arithmetic Transformers

Samy Jelassi, Stéphane d'Ascoli, Carles Domingo-Enrich +3

We examine how transformers cope with two challenges: learning basic integer arithmetic, and generalizing to longer sequences than seen during training. We find that relative posit…