67 citations · 126 across the 12 of their papers we have counts for
13 papers · 1 filter
Parameter-Efficient Transfer Learning under Federated Learning for Automatic Speech Recognition
Xuan Kan, Yonghui Xiao, Tien-Ju Yang +2
This work explores the challenge of enhancing Automatic Speech Recognition (ASR) model performance across various user-specific domains while preserving user data privacy. We emplo…
Efficient Adapters for Giant Speech Models
Nanxin Chen, Izhak Shafran, Yu Zhang +4
Large pre-trained speech models are widely used as the de-facto paradigm, especially in scenarios when there is a limited amount of labeled data available. However, finetuning all…
A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation
Yosuke Higuchi, Nanxin Chen, Yuya Fujita +6
Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to aut…
WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis
Nanxin Chen, Yu Zhang, Heiga Zen +4
This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density…
Focus on the present: a regularization method for the ASR source-target attention layer
Nanxin Chen, Piotr Żelasko, Jesús Villalba +1
This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classificati…
WaveGrad: Estimating Gradients for Waveform Generation
Nanxin Chen, Yu Zhang, Heiga Zen +3
This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and di…