Attention based end to end Speech Recognition for Voice Search in Hindi and English
arXiv:2111.10208 · doi:10.1145/3503162.3503173
Abstract
We describe here our work with automatic speech recognition (ASR) in the context of voice search functionality on the Flipkart e-Commerce platform. Starting with the deep learning architecture of Listen-Attend-Spell (LAS), we build upon and expand the model design and attention mechanisms to incorporate innovative approaches including multi-objective training, multi-pass training, and external rescoring using language models and phoneme based losses. We report a relative WER improvement of 15.7% on top of state-of-the-art LAS models using these modifications. Overall, we report an improvement of 36.9% over the phoneme-CTC system. The paper also provides an overview of different components that can be tuned in a LAS-based system.
Accepted at Forum for Information Retrieval Evaluation (FIRE) 2021
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- An Overview of Multi-Task Learning in Deep Neural Networks
- Attention-Based Models for Speech Recognition
- Sequence Transduction with Recurrent Neural Networks
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Transfer Learning Approaches for Streaming End-to-End Speech Recognition System