activity
20172022
most citedSpeaker Adaptation for Attention-Based End-to-End Speech Recognition

41 citations · 101 across the 21 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20251 cited

Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study

Menglong Cui, Pengzhi Gao, Wei Liu +2

Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. I…

cs.CL2022

Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition

Yashesh Gaur, Nick Kibre, Jian Xue +5

Automatic Speech Recognition (ASR) systems typically yield output in lexical form. However, humans prefer a written form output. To bridge this gap, ASR systems usually employ Inve…

cs.CL20201 cited

Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR

Hirofumi Inaguma, Yashesh Gaur, Liang Lu +2

Recently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. Howe…

cs.CL2020

Serialized Output Training for End-to-End Overlapped Speech Recognition

Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang +2

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instea…

cs.CL201941 cited

Speaker Adaptation for Attention-Based End-to-End Speech Recognition

Zhong Meng, Yashesh Gaur, Jinyu Li +1

We propose three regularization-based speaker adaptation approaches to adapt the attention-based encoder-decoder (AED) model with very limited adaptation data from target speakers…

cs.CL2017

Robust Speech Recognition Using Generative Adversarial Networks

Anuroop Sriram, Heewoo Jun, Yashesh Gaur +1

This paper describes a general, scalable, end-to-end framework that uses the generative adversarial network (GAN) objective to enable robust speech recognition. Encoders trained wi…