25 citations · 63 across the 13 of their papers we have counts for
5 papers
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Tianrui Wang, Long Zhou, Ziqiang Zhang +6
Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we propose V…
Code-Switching Text Generation and Injection in Mandarin-English ASR
Haibin Yu, Yuxuan Hu, Yao Qian +7
Code-switching speech refers to a means of expression by mixing two or more languages within a single utterance. Automatic Speech Recognition (ASR) with End-to-End (E2E) modeling f…
Target Sound Extraction with Variable Cross-modality Clues
Chenda Li, Yao Qian, Zhuo Chen +5
Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture o…
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang +10
We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec lan…
Self-Supervised Learning for speech recognition with Intermediate layer supervision
Chengyi Wang, Yu Wu, Sanyuan Chen +4
Recently, pioneer work finds that speech pre-trained models can solve full-stack speech processing tasks, because the model utilizes bottom layers to learn speaker-related informat…