12 citations · 16 across the 3 of their papers we have counts for
3 papers
VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition
Jinhan Wang, Xiaosu Tong, Jinxi Guo +2
While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed…
Enhancing ASR for Stuttered Speech with Limited Data Using Detect and Pass
Olabanji Shonibare, Xiaosu Tong, Venkatesh Ravichandran
It is estimated that around 70 million people worldwide are affected by a speech disorder called stuttering. With recent advances in Automatic Speech Recognition (ASR), voice assis…
Streaming ResLSTM with Causal Mean Aggregation for Device-Directed Utterance Detection
Xiaosu Tong, Che-Wei Huang, Sri Harish Mallidi +5
In this paper, we propose a streaming model to distinguish voice queries intended for a smart-home device from background speech. The proposed model consists of multiple CNN layers…