2 papers
cs.LG2026
From to : Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
Sizhe Dang, Jiaqi Shao, Xiaodong Zheng +3
As foundation models continue to scale, pretraining increasingly relies on data-parallel distributed optimization, making bandwidth-limited gradient synchronization a key bottlenec…
cs.CL2025
Frustratingly Easy Task-aware Pruning for Large Language Models
Yuanhe Tian, Junjie Liu, Xican Yang +2
Pruning provides a practical solution to reduce the resources required to run large language models (LLMs) to benefit from their effective capabilities as well as control their cos…