4 papers
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
Jiatong Shi, Chunlei Zhang, Jinchuan Tian +4
Recent advances in speech language models (LLMs) have extended textual LLMs to the speech domain, but balancing speech understanding and generation remains challenging, especially…
Towards Unsupervised Speech Recognition at the Syllable-Level
Liming Wang, Junrui Ni, Kai-Wei Chang +4
Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in…
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
Kaizhi Qian, Xulin Fan, Junrui Ni +4
Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to cap…
Towards Unsupervised Speech Recognition Without Pronunciation Models
Junrui Ni, Liming Wang, Yang Zhang +4
Recent advancements in supervised automatic speech recognition (ASR) have achieved remarkable performance, largely due to the growing availability of large transcribed speech corpo…