3 papers
cs.LG2026
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
Ke Li, Dong An, Xiaoling Zang +6
Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activations contain outliers, but tha…
eess.AS2024
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
Niko Moritz, Ruiming Xie, Yashesh Gaur +5
We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recogn…
eess.AS2024
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Wonjune Kang, Junteng Jia, Chunyang Wu +8
This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system w…