2 papers
cs.LG2026
Sliced-Wasserstein Distribution Alignment Loss Improves the Ultra-Low-Bit Quantization of Large Language Models
Deyu Cao, Yixin Yin, Samin Aref
The benefits of most large language models come with steep and often hidden economic and environmental costs due to their resource usage inefficiency during deployment. Model quant…
cs.LG2025
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
Deyu Cao, Samin Aref
The growing use of large language models has raised environmental and economic concerns about their intensity of resource usage during inference. Serving these models to each user…