2 papers
cs.LG2025
DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
Tianteng Gu, Bei Liu, Bo Xiao +3
Pruning is a widely used technique to compress large language models (LLMs) by removing unimportant weights, but it often suffers from significant performance degradation - especia…
cs.DC2025
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
Xinyuan Lin, Chenlu Li, Zongle Huang +5
Larger model sizes and longer sequence lengths have empowered the Large Language Model (LLM) to achieve outstanding performance across various domains. However, this progress bring…