most citedELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

2 citations · 5 across the 9 of their papers we have counts for

collaborators

9 papers

eess.AS2024

The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge

Yiwei Guo, Chenrun Wang, Yifan Yang +9

Discrete speech tokens have been more and more popular in multiple speech processing fields, including automatic speech recognition (ASR), text-to-speech (TTS) and singing voice sy…

cs.SD2024

SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

Junjie Li, Yiwei Guo, Xie Chen +1

Zero-shot voice conversion (VC) aims to transfer the source speaker timbre to arbitrary unseen target speaker timbre, while keeping the linguistic content unchanged. Although the v…

cs.CL20242 cited

ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Yakun Song, Zhuo Chen, Xiaofei Wang +2

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, exi…

eess.AS20241 cited

EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Wenxi Chen, Yuzhe Liang, Ziyang Ma +2

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational…

eess.AS20231 cited

Expressive TTS Driven by Natural Language Prompts Using Few Human Annotations

Hanglei Zhang, Yiwei Guo, Sen Liu +2

Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the…

eess.AS2023

Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition

Zhisheng Zheng, Ziyang Ma, Yu Wang +1

In recent years, speech-based self-supervised learning (SSL) has made significant progress in various tasks, including automatic speech recognition (ASR). An ASR model with decent…