1 paper
Yanyi Li, Yimu Zhang, Cong Fang
Activations have become the primary memory bottleneck in large-batch LLM training. However, existing compression methods fail to exploit the spectral structure of activations, resu…