activity
20222026
most citedLeveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation

2 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

Controllable Accent Normalization via Discrete Diffusion

Qibing Bai, Yuhan Du, Tom Ko +3

Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retent…

eess.AS2026

CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data

Qibing Bai, Shuhao Shi, Shuai Wang +3

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. I…

eess.AS2025

Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data

Qibing Bai, Sho Inoue, Shuai Wang +3

Accent normalization converts foreign-accented speech into native-like speech while preserving speaker identity. We propose a novel pipeline using self-supervised discrete tokens a…

eess.AS2024

Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Zhijun Liu, Shuai Wang, Sho Inoue +2

Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete s…

eess.AS2023

Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition

Shuai Wang, Qibing Bai, Qi Liu +5

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage lar…