activity
20212026
most citedSynthetic Cross-accent Data Augmentation for Automatic Speech Recognition

3 citations · 6 across the 10 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

GSRM: Generative Speech Reward Model for Speech RLHF

Maohao Shen, Tejas Jayashankar, Osama Hanna +10

Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic natura…

cs.SD2022★ 1 cited

Voice-preserving Zero-shot Multiple Accent Conversion

Mumin Jin, Prashant Serai, Jilong Wu +3

Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker's accent. For native speakers, understand…

cs.SD2022

Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders

Jason Fong, Yun Wang, Prabhav Agrawal +4

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edi…

cs.SD2021

VocBench: A Neural Vocoder Benchmark for Speech Synthesis

Ehab A. AlBadawy, Andrew Gibiansky, Qing He +3

Neural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on sy…

cs.SD2021★ 2 cited

Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling

Qing He, Zhiping Xiu, Thilo Koehler +1

Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…