activity
20172022
most citedStream Attention for far-field multi-microphone ASR

2 citations · 3 across the 7 of their papers we have counts for

collaborators

11 papers

eess.AS2022

Blind Signal Dereverberation for Machine Speech Recognition

Samik Sadhu, Hynek Hermansky

We present a method to remove unknown convolutive noise introduced to speech by reverberations of recording environments, utilizing some amount of training speech data from the rev…

cs.SD2022

Complex Frequency Domain Linear Prediction: A Tool to Compute Modulation Spectrum of Speech

Samik Sadhu, Hynek Hermansky

Conventional Frequency Domain Linear Prediction (FDLP) technique models the squared Hilbert envelope of speech with varied degrees of approximation which can be sampled at the requ…

eess.AS2021

Radically Old Way of Computing Spectra: Applications in End-to-End ASR

Samik Sadhu, Hynek Hermansky

We propose a technique to compute spectrograms using Frequency Domain Linear Prediction (FDLP) that uses all-pole models to fit the squared Hilbert envelope of speech in different…

cs.SD2021

Two-Stage Augmentation and Adaptive CTC Fusion for Improved Robustness of Multi-Stream End-to-End ASR

Ruizhi Li, Gregory Sell, Hynek Hermansky

Performance degradation of an Automatic Speech Recognition (ASR) system is commonly observed when the test acoustic condition is different from training. Hence, it is essential to…

cs.CL2019

A practical two-stage training strategy for multi-stream end-to-end speech recognition

Ruizhi Li, Gregory Sell, Xiaofei Wang +2

The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study o…

eess.AS2019

Multi-Stream End-to-End Speech Recognition

Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi +3

Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The…