activity
20172023
most citedConditional Teacher-Student Learning

108 citations · 427 across the 46 of their papers we have counts for

collaborators
Showing eess.ASShow all

37 papers · 1 filter

eess.AS2023

SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

Xiaofei Wang, Manthan Thakker, Zhuo Chen +7

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models…

eess.AS2023

Pre-training End-to-end ASR Models with Augmented Speech Samples Queried by Text

Eric Sun, Jinyu Li, Jian Xue +1

In end-to-end automatic speech recognition system, one of the difficulties for language expansion is the limited paired speech and text training data. In this paper, we propose a n…

eess.AS2023

On decoder-only architecture for speech-to-text and large language model integration

Jian Wu, Yashesh Gaur, Zhuo Chen +8

Large language models (LLMs) have achieved remarkable success in the field of natural language processing, enabling better human-computer interaction using natural language. Howeve…

eess.AS2022

Speech separation with large-scale self-supervised learning

Zhuo Chen, Naoyuki Kanda, Jian Wu +6

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the ex…

eess.AS20221 cited

Simulating realistic speech overlaps improves multi-talker ASR

Muqiao Yang, Naoyuki Kanda, Xiaofei Wang +5

Multi-talker automatic speech recognition (ASR) has been studied to generate transcriptions of natural conversation including overlapping speech of multiple speakers. Due to the di…

eess.AS2022

Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition

Zili Huang, Zhuo Chen, Naoyuki Kanda +6

Self-supervised learning (SSL), which utilizes the input data itself for representation learning, has achieved state-of-the-art results for various downstream speech tasks. However…