7 papers · 1 filter
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
Pengcheng Li, Jianzong Wang, Xulong Zhang +3
One-shot voice conversion aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing methods face difficulties…
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
Haoxiang Shi, Jianzong Wang, Xulong Zhang +3
Although current Text-To-Speech (TTS) models are able to generate high-quality speech samples, there are still challenges in developing emotion intensity controllable TTS. Most exi…
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
Yimin Deng, Jianzong Wang, Xulong Zhang +2
Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are…
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
Jianzong Wang, Ziqi Liang, Xulong Zhang +2
In recent years, Transformer networks have shown remarkable performance in speech recognition tasks. However, their deployment poses challenges due to high computational and storag…
EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning
Ziqi Liang, Jianzong Wang, Xulong Zhang +3
Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into a…
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
Jianzong Wang, Pengcheng Li, Xulong Zhang +2
Singing voice beautifying is a novel task that has application value in people's daily life, aiming to correct the pitch of the singing voice and improve the expressiveness without…