activity
20172023
most citedConformer: Convolution-augmented Transformer for Speech Recognition

387 citations · 711 across the 19 of their papers we have counts for

collaborators

34 papers

eess.AS20234 cited

Efficient Adapters for Giant Speech Models

Nanxin Chen, Izhak Shafran, Yu Zhang +4

Large pre-trained speech models are widely used as the de-facto paradigm, especially in scenarios when there is a limited amount of labeled data available. However, finetuning all…

cs.CL20221 cited

Textless Direct Speech-to-Speech Translation with Discrete Speech Representation

Xinjian Li, Ye Jia, Chung-Cheng Chiu

Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade sys…

eess.AS202212 cited

Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data

Alëna Aksënova, Zhehuai Chen, Chung-Cheng Chiu +8

Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work fo…

eess.AS20211 cited

Cross-attention conformer for context modeling in speech enhancement for ASR

Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley +2

This work introduces \emph{cross-attention conformer}, an attention-based architecture for context modeling in speech enhancement. Given that the context information can often be s…

cs.LG202115 cited

W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Yu-An Chung, Yu Zhang, Wei Han +4

Motivated by the success of masked language modeling~(MLM) in pre-training natural language processing models, we propose w2v-BERT that explores MLM for self-supervised speech repr…

cs.CL2021

Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models

Thibault Doutre, Wei Han, Chung-Cheng Chiu +3

Streaming end-to-end automatic speech recognition (ASR) systems are widely used in everyday applications that require transcribing speech to text in real-time. Their minimal latenc…