6 papers
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling
Haolong Zheng, Yuanzhuo Hu, Xinyu Liang +7
CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and l…
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
Xulin Fan, Vishal Sunder, Samuel Thomas +3
Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer…
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
George Saon, Samuel Thomas, Takashi Fukuda +3
We propose self-speculative decoding for speech-aware LLMs by using the CTC encoder as a draft model to accelerate auto-regressive (AR) inference and improve ASR accuracy. Our thre…
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
Avihu Dekel, Samuel Thomas, Takashi Fukada +1
While autoregressive (AR) LLM-based ASR systems achieve strong accuracy, their sequential decoding limits parallelism and incurs high latency. We propose NLE, a non-autoregressive…
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
George Saon, Avihu Dekel, Alexander Brooks +21
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…
A Non-autoregressive Model for Joint STT and TTS
Vishal Sunder, Brian Kingsbury, George Saon +5
In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimoda…