5 papers
Joint Speech Recognition and Audio Captioning
Chaitanya Narisetty, Emiru Tsunoo, Xuankai Chang +3
Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remo…
Run-and-back stitch search: novel block synchronous decoding for streaming encoder-decoder ASR
Emiru Tsunoo, Chaitanya Narisetty, Michael Hentschel +2
A streaming style inference of encoder-decoder automatic speech recognition (ASR) system is important for reducing latency, which is essential for interactive use cases. To this en…
Data Augmentation Methods for End-to-end Speech Recognition on Distant-Talk Scenarios
Emiru Tsunoo, Kentaro Shibata, Chaitanya Narisetty +2
Although end-to-end automatic speech recognition (E2E ASR) has achieved great performance in tasks that have numerous paired data, it is still challenging to make E2E ASR robust ag…
Bayesian Non-Parametric Multi-Source Modelling Based Determined Blind Source Separation
Chaitanya Narisetty, Tatsuya Komatsu, Reishi Kondo
This paper proposes a determined blind source separation method using Bayesian non-parametric modelling of sources. Conventionally source signals are separated from a given set of…
Modelling of Sound Events with Hidden Imbalances Based on Clustering and Separate Sub-Dictionary Learning
Chaitanya Narisetty, Tatsuya Komatsu, Reishi Kondo
This paper proposes an effective modelling of sound event spectra with a hidden data-size-imbalance, for improved Acoustic Event Detection (AED). The proposed method models each ev…