5 papers
Interactive In-Meeting Speaker Correction with Human Feedback
Xinlu He, Yiwen Guan, Badrivishal Paurana +3
Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows can potentially enable higher accu…
Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation
Yiwen Guan, Jacob Whitehill
Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In addition, translation quality can suffer for…
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
Yiwen Guan, Viet Anh Trinh, Vivek Voleti +1
Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior…
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
Yiwen Guan, Viet Anh Trinh, Vivek Voleti +1
Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities imp…
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
Viet Anh Trinh, Rosy Southwell, Yiwen Guan +3
Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, spe…