2 papers
cs.SD2025
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
Yu Chen, Hongxu Zhu, Jiadong Wang +2
Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially…
eess.AS2025
Exploring Length Generalization For Transformer-based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian +2
Transformer network architecture has proven effective in speech enhancement. However, as its core module, self-attention suffers from quadratic complexity, making it infeasible for…