activity
20222024
most citedAn Embarrassingly Simple Approach for LLM with Strong ASR Capacity

4 citations · 4 across the 4 of their papers we have counts for

collaborators

6 papers

cs.SD20244 cited

CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Zhihao Du, Qian Chen, Shiliang Zhang +9

Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity. In…

physics.soc-ph202438 cited

Intercity Connectivity and Innovation

Xiaofan Liang, César A. Hidalgo, Pierre-Alexandre Balland +2

Urban outputs, from economy to innovation, are known to grow as a power of a city's population. But, since large cities tend to be central in transportation and communication netwo…

cs.CL20244 cited

An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Ziyang Ma, Guanrou Yang, Yifan Yang +8

In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and…

cs.SD2023

FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec

Zhihao Du, Shiliang Zhang, Kai Hu +1

This paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible t…

cs.CL2023

Improving BERT with Hybrid Pooling Network and Drop Mask

Qian Chen, Wen Wang, Qinglin Zhang +3

Transformer-based pre-trained language models, such as BERT, achieve great success in various natural language understanding tasks. Prior research found that BERT captures a rich h…

cs.SD2022

PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification

Siqi Zheng, Hongbin Suo, Qian Chen

Speaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as f…