17 citations · 22 across the 2 of their papers we have counts for
4 papers
Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset
Xie Chen, Yu Wu, Zhenghao Wang +2
Recently, Transformer based end-to-end models have achieved great success in many areas including speech recognition. However, compared to LSTM models, the heavy computational cost…
Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
Jinyu Li, Rui Zhao, Zhong Meng +8
Because of its streaming nature, recurrent neural network transducer (RNN-T) is a very promising end-to-end (E2E) model that may replace the popular hybrid model for automatic spee…
Advances in Online Audio-Visual Meeting Transcription
Takuya Yoshioka, Igor Abramovski, Cem Aksoylar +23
This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its abilit…
Cracking the cocktail party problem by multi-beam deep attractor network
Zhuo Chen, Jinyu Li, Xiong Xiao +4
While recent progresses in neural network approaches to single-channel speech separation, or more generally the cocktail party problem, achieved significant improvement, their perf…