3 papers
eess.AS2024
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Yifan Yang, Shujie Liu, Jinyu Li +10
This paper introduces Interleaved Speech-Text Language Model (IST-LM) for zero-shot streaming Text-to-Speech (TTS). Unlike many previous approaches, IST-LM is directly trained on i…
eess.AS2024
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
Ruchao Fan, Bo Ren, Yuxuan Hu +3
Integrating speech into LLM (speech-LLM) has gaining increased attention recently. The mainstream solution is to connect a well-trained speech encoder and LLM with a neural adapter…
cs.CL2024
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
Rui Zhao, Jinyu Li, Ruchao Fan +1
Models for streaming speech translation (ST) can achieve high accuracy and low latency if they're developed with vast amounts of paired audio in the source language and written tex…