4 papers
Latent Speech-Text Transformer
Yen-Ju Lu, Yashesh Gaur, Wei Zhou +8
Auto-regressive speech-text models pre-trained on interleaved text tokens and discretized speech tokens demonstrate strong speech understanding and generation, yet remain substanti…
Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velazquez +2
We present Paired by the Teacher (PbT), a two-stage teacher-student pipeline that synthesizes accurate input-output pairs without human labels or parallel data. In many low-resourc…
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
Thomas Thebaud, Yen-Ju Lu, Matthew Wiesner +2
In dialogue transcription pipelines, Large Language Models (LLMs) are frequently employed in post-processing to improve grammar, punctuation, and readability. We explore a compleme…
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
Helin Wang, Jiarui Hai, Yen-Ju Lu +3
In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing t…