most citedAn Embarrassingly Simple Approach for LLM with Strong ASR Capacity

4 citations · 13 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SD20244 cited

CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Zhihao Du, Qian Chen, Shiliang Zhang +9

Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity. In…

cs.CL20244 cited

An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Ziyang Ma, Guanrou Yang, Yifan Yang +8

In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and…

cs.SD2023

SA-Paraformer: Non-autoregressive End-to-End Speaker-Attributed ASR

Yangze Li, Fan Yu, Yuhao Liang +5

Joint modeling of multi-speaker ASR and speaker diarization has recently shown promising results in speaker-attributed automatic speech recognition (SA-ASR).Although being able to…

cs.SD2023

FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec

Zhihao Du, Shiliang Zhang, Kai Hu +1

This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible t…

cs.SD20231 cited

The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR

Yuhao Liang, Mohan Shi, Fan Yu +11

With the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle…

eess.AS2023

CASA-ASR: Context-Aware Speaker-Attributed ASR

Mohan Shi, Zhihao Du, Qian Chen +5

Recently, speaker-attributed automatic speech recognition (SA-ASR) has attracted a wide attention, which aims at answering the question ``who spoke what''. Different from modular s…