1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.SD2025
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
Haoran Zhou, Xingchen Song, Brendan Fahy +9
OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence…
eess.AS2024★ 1 cited
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
Qiquan Zhang, Meng Ge, Hongxu Zhu +4
Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to…
cs.SD2023
Ripple sparse self-attention for monaural speech enhancement
Qiquan Zhang, Hongxu Zhu, Qi Song +3
The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally…