most citedDiscrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

eess.AS2025

Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs

Xinlu He, Swayambhu Nath Ray, Harish Mallidi +5

Unified architectures in multimodal large language models (MLLM) have shown promise in handling diverse tasks within a single framework. In the text-to-speech (TTS) task, current M…

cs.CL2025

Interactive In-Meeting Speaker Correction with Human Feedback

Xinlu He, Yiwen Guan, Badrivishal Paurana +3

Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows can potentially enable higher accu…

cs.CL2025

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context

Viet Anh Trinh, Xinlu He, Jacob Whitehill

Classroom speech and lectures often contain named entities (NEs) such as names of people and special terminology. While automatic speech recognition (ASR) systems have achieved rem…

cs.CL2025

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

Xinlu He, Jacob Whitehill

Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual s…

cs.CL20241 cited

Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

Viet Anh Trinh, Rosy Southwell, Yiwen Guan +3

Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, spe…