2 papers
cs.CV2025
Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
Zekang Zheng, Haokun Li, Yaofo Chen +2
Model quantization reduces neural network parameter precision to achieve compression, but often compromises accuracy. Existing post-training quantization (PTQ) methods employ itera…
cs.CL2025
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
Jinwu Hu, Wei Zhang, Yufeng Wang +4
Large Language Models (LLMs) have shown outstanding performance across a variety of tasks, partly due to advanced prompting techniques. However, these techniques often require leng…