2 papers
cs.LG2025
TensorGRaD: Tensor Gradient Robust Decomposition for Memory-Efficient Neural Operator Training
Sebastian Loeschcke, David Pitt, Robert Joseph George +5
Scientific problems require resolving multi-scale phenomena across different resolutions and learning solution operators in infinite-dimensional function spaces. Neural operators p…
cs.LG2025
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
DiJia Su, Andrew Gu, Jane Xu +2
Large language models (LLMs) have revolutionized natural language understanding and generation but face significant memory bottlenecks during training. GaLore, Gradient Low-Rank Pr…