1 citations · 1 across the 3 of their papers we have counts for
5 papers · 1 filter
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
Sung-Lin Yeh, Wei Zhou, Gil Keren +6
Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downstream objectives, the extracte…
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Wonjune Kang, Junteng Jia, Chunyang Wu +8
This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system w…
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
Wei Zhou, Junteng Jia, Leda Sari +2
CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications. In this work, w…
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
Yufeng Yang, Desh Raj, Ju Lin +8
The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. Howeve…
Faster Speech-LLaMA Inference with Multi-token Prediction
Desh Raj, Gil Keren, Junteng Jia +2
Large language models (LLMs) have become proficient at solving a wide variety of tasks, including those involving multi-modal inputs. In particular, instantiating an LLM (such as L…