6 papers
GSRM: Generative Speech Reward Model for Speech RLHF
Maohao Shen, Tejas Jayashankar, Osama Hanna +10
Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…
VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
Yancheng Wang, Osama Hanna, Ruiming Xie +11
Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features suc…
Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT
Zeeshan Ahmed, Frank Seide, Niko Moritz +5
This paper tackles several challenges that arise when integrating Automatic Speech Recognition (ASR) and Machine Translation (MT) for real-time, on-device streaming speech translat…
Non-Monotonic Attention-based Read/Write Policy Learning for Simultaneous Translation
Zeeshan Ahmed, Frank Seide, Zhe Liu +6
Simultaneous or streaming machine translation generates translation while reading the input stream. These systems face a quality/latency trade-off, aiming to achieve high translati…
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
Niko Moritz, Ruiming Xie, Yashesh Gaur +5
We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recogn…
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
Jinzheng Zhao, Niko Moritz, Egor Lakomkin +7
Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumu…