4 papers · 1 filter
ZO-SAM: Zero-Order Sharpness-Aware Minimization for Efficient Sparse Training
Jie Ji, Gen Li, Kaiyuan Deng +2
Deep learning models, despite their impressive achievements, suffer from high computational costs and memory requirements, limiting their usability in resource-constrained environm…
From Bits to Chips: An LLM-based Hardware-Aware Quantization Agent for Streamlined Deployment of LLMs
Kaiyuan Deng, Hangyu Zheng, Minghai Qing +11
Deploying models, especially large language models (LLMs), is becoming increasingly attractive to a broader user base, including those without specialized expertise. However, due t…
The Right to be Forgotten in Pruning: Unveil Machine Unlearning on Sparse Models
Yang Xiao, Gen Li, Jie Ji +3
Machine unlearning aims to efficiently eliminate the memory about deleted data from trained models and address the right to be forgotten. Despite the success of existing unlearning…
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
Mingyu Cao, Gen Li, Jie Ji +6
Mixture-of-Experts (MoE) has garnered significant attention for its ability to scale up neural networks while utilizing the same or even fewer active parameters. However, MoE does…