1 citations · 1 across the 4 of their papers we have counts for
4 papers
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
Yu Chen, Hongxu Zhu, Jiadong Wang +2
Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially…
Exploring Length Generalization For Transformer-based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian +2
Transformer network architecture has proven effective in speech enhancement. However, as its core module, self-attention suffers from quadratic complexity, making it infeasible for…
An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian +2
The use of Transformer architectures has facilitated remarkable progress in speech enhancement. Training Transformers using substantially long speech utterances is often infeasible…
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
Qiquan Zhang, Meng Ge, Hongxu Zhu +4
Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to…