3 papers
cs.CL2026
FormalASR: End-to-End Spoken Chinese to Formal Text
Wanyi Ning, Yinshang Guo, Haitao Qian +3
Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informal spoken structures that are o…
cs.CL2026
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
Hao Zhang, Zhibin Zhang, Guangxin Wu +3
Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinder deployment in resource-con…
cs.LG2025
MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
Jinguang Wang, Jingyu Wang, Haifeng Sun +6
Quantization has been widely used to compress and accelerate inference of large language models (LLMs). Existing methods focus on exploring the per-token dynamic calibration to ens…