3 papers
cs.CL2019
A practical two-stage training strategy for multi-stream end-to-end speech recognition
Ruizhi Li, Gregory Sell, Xiaofei Wang +2
The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study o…
cs.CL2019
Performance Monitoring for End-to-End Speech Recognition
Ruizhi Li, Gregory Sell, Hynek Hermansky
Measuring performance of an automatic speech recognition (ASR) system without ground-truth could be beneficial in many scenarios, especially with data from unseen domains, where pe…
eess.AS2018
Joint Acoustic and Class Inference for Weakly Supervised Sound Event Detection
Sandeep Kothinti, Keisuke Imoto, Debmalya Chakrabarty +3
Sound event detection is a challenging task, especially for scenes with multiple simultaneous events. While event classification methods tend to be fairly accurate, event localizat…