Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment
Sieun Hyeon, Jaeyoung Do
Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. We argue that this view i…
cs.LG2026
Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
Sieun Hyeon, Jaeyoung Do
Mixture-of-Experts (MoE) models scale capacity efficiently, but their massive parameter footprint creates a deployment-time memory bottleneck. We organize retraining-free MoE compr…