3 citations · 6 across the 10 of their papers we have counts for
5 papers · 1 filter
GSRM: Generative Speech Reward Model for Speech RLHF
Maohao Shen, Tejas Jayashankar, Osama Hanna +10
Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…
Voice-preserving Zero-shot Multiple Accent Conversion
Mumin Jin, Prashant Serai, Jilong Wu +3
Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker's accent. For native speakers, understand…
Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders
Jason Fong, Yun Wang, Prabhav Agrawal +4
Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edi…
VocBench: A Neural Vocoder Benchmark for Speech Synthesis
Ehab A. AlBadawy, Andrew Gibiansky, Qing He +3
Neural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on sy…
Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
Qing He, Zhiping Xiu, Thilo Koehler +1
Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…