Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SignMuon: Communication-Efficient Distributed Muon Optimization
Neel Mishra, Kushagara Trivedi, Pawan Kumar
Distributed training of large neural networks is bottlenecked by full-precision gradient communication and by coordinatewise optimizers that ignore the matrix structure of weight t…
cs.LG2025
Hierarchical Sparse Plus Low Rank Compression of LLM
Pawan Kumar, Aditi Gupta
Modern large language models (LLMs) place extraordinary pressure on memory and compute budgets, making principled compression indispensable for both deployment and continued traini…