activity
20192022
most citedOn End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments

49 citations · 68 across the 6 of their papers we have counts for

collaborators

8 papers

eess.AS202213 cited

Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition

Catalin Zorila, Rama Doddipatla

Improving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they ty…

eess.AS2022

Transformer-based Streaming ASR with Cumulative Attention

Mohan Li, Shucong Zhang, Catalin Zorila +1

In this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monoton…

cs.SD20212 cited

Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation

Jisi Zhang, Catalin Zorila, Rama Doddipatla +1

In this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixt…

eess.AS20211 cited

Head-synchronous Decoding for Transformer-based Streaming ASR

Mohan Li, Catalin Zorila, Rama Doddipatla

Online Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder…

eess.AS2021

Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism

Jisi Zhang, Catalin Zorila, Rama Doddipatla +1

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environ…

eess.AS20203 cited

Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps

Mohan Li, Catalin Zorila, Rama Doddipatla

Transformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent struct…