most citedELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

2 citations · 5 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2024

The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge

Yiwei Guo, Chenrun Wang, Yifan Yang +9

Discrete speech tokens have been more and more popular in multiple speech processing fields, including automatic speech recognition (ASR), text-to-speech (TTS) and singing voice sy…

eess.AS20241 cited

EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

Wenxi Chen, Yuzhe Liang, Ziyang Ma +2

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational…

eess.AS20231 cited

Expressive TTS Driven by Natural Language Prompts Using Few Human Annotations

Hanglei Zhang, Yiwei Guo, Sen Liu +2

Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the…

eess.AS2023

Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition

Zhisheng Zheng, Ziyang Ma, Yu Wang +1

In recent years, speech-based self-supervised learning (SSL) has made significant progress in various tasks, including automatic speech recognition (ASR). An ASR model with decent…

eess.AS2023

Blank-regularized CTC for Frame Skipping in Neural Transducer

Yifan Yang, Xiaoyu Yang, Liyong Guo +6

Neural Transducer and connectionist temporal classification (CTC) are popular end-to-end automatic speech recognition systems. Due to their frame-synchronous design, blank symbols…

eess.AS2023

Front-End Adapter: Adapting Front-End Input of Speech based Self-Supervised Learning for Speech Recognition

Xie Chen, Ziyang Ma, Changli Tang +2

Recent years have witnessed a boom in self-supervised learning (SSL) in various areas including speech processing. Speech based SSL models present promising performance in a range…