1 paper
Dacheng Li, Rulin Shao, Anze Xie +5
FlashAttention (Dao, 2023) effectively reduces the quadratic peak memory usage to linear in training transformer-based large language models (LLMs) on a single GPU. In this paper,…