2 papers
cs.CV2025
Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
Zekang Zheng, Haokun Li, Yaofo Chen +2
Model quantization reduces neural network parameter precision to achieve compression, but often compromises accuracy. Existing post-training quantization (PTQ) methods employ itera…
cs.CL2025
Core Context Aware Transformers for Long Context Language Modeling
Yaofo Chen, Zeng You, Shuhai Zhang +4
Transformer-based Large Language Models (LLMs) have exhibited remarkable success in extensive tasks primarily attributed to self-attention mechanism, which requires a token to cons…