1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
DiJia Su, Andrew Gu, Jane Xu +2
Large language models (LLMs) have revolutionized natural language understanding and generation but face significant memory bottlenecks during training. GaLore, Gradient Low-Rank Pr…