most citedSpeak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

25 citations · 51 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD20231 cited

The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR

Yuhao Liang, Mohan Shi, Fan Yu +11

With the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle…

eess.AS20231 cited

Adapting Multi-Lingual ASR Models for Handling Multiple Talkers

Chenda Li, Yao Qian, Zhuo Chen +5

State-of-the-art large-scale universal speech models (USMs) show a decent automatic speech recognition (ASR) performance across multiple domains and languages. However, it remains…

cs.CL202317 cited

VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Tianrui Wang, Long Zhou, Ziqiang Zhang +6

Recent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we propose V…

cs.CL202325 cited

Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Ziqiang Zhang, Long Zhou, Chengyi Wang +10

We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec lan…

eess.AS2023

Speaker Change Detection for Transformer Transducer ASR

Jian Wu, Zhuo Chen, Min Hu +2

Speaker change detection (SCD) is an important feature that improves the readability of the recognized words from an automatic speech recognition (ASR) system by breaking the word…

cs.CL20177 cited

End-to-End Attention based Text-Dependent Speaker Verification

Shi-Xiong Zhang, Zhuo Chen, Yong Zhao +2

A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetically discriminative/speaker discriminative DNNs as…