2 papers
cs.SD2026
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
Xinmeng Xu, Haoran Xie, S. Joe Qin +3
Stage-wise audio-visual encoders propagate fused intermediate states across layers, making the formation of later representations depend on the readiness of earlier fusion states.…
cs.SD2025
Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
Jizhen Li, Weiping Tu, Yuhong Yang +3
Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to subs…