49 citations · 68 across the 6 of their papers we have counts for
8 papers
Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition
Catalin Zorila, Rama Doddipatla
Improving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they ty…
Transformer-based Streaming ASR with Cumulative Attention
Mohan Li, Shucong Zhang, Catalin Zorila +1
In this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monoton…
Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation
Jisi Zhang, Catalin Zorila, Rama Doddipatla +1
In this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixt…
Head-synchronous Decoding for Transformer-based Streaming ASR
Mohan Li, Catalin Zorila, Rama Doddipatla
Online Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder…
Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism
Jisi Zhang, Catalin Zorila, Rama Doddipatla +1
In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environ…
Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
Mohan Li, Catalin Zorila, Rama Doddipatla
Transformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent struct…