5 papers
SB-RF: Schrödinger Bridge Rectified Flow for One-Step Robust Speech Enhancement
Caixia Lu, Xueyang Lv, Penglong Hu +1
Generative models have shown promising results for speech enhancement (SE), but they often rely on multi-step inference, limiting low-latency deployment. We propose SB-RF, a one-st…
SHB-AE: Spherical harmonic beamforming based Ambisonics encoding and upscaling method for smartphone microphone array
Yuhuan You, Yufan Qian, Tianshu Qu +2
With the rapid development of virtual reality (VR) and augmented reality (AR), spatial audio recording and reproduction have gained increasing research interest. Higher Order Ambis…
Flow-HOA: Generative Joint Optimization for Ambisonics Encoding via Flow Matching
Yuhuan You, Yufan Qian, Tianshu Qu +2
Higher-Order Ambisonics (HOA) encoding from sparse, irregular microphone arrays remains a critical challenge for consumer spatial audio capture in immersive communication and XR. W…
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
Jiajun Yuan, Xiaochen Wang, Yuhang Xiao +3
Speech super-resolution (SR) reconstructs high-fidelity wideband speech from low-resolution inputs-a task that necessitates reconciling global harmonic coherence with local transie…
CogSR: Semantic-Aware Speech Super-Resolution via Chain-of-Thought Guided Flow Matching
Jiajun Yuan, Xiaochen Wang, Yuhang Xiao +3
Applying speech super-resolution (SR) to recordings with severely low sampling rates is a critical challenge in digital archiving and investigative audio recovery. In these scenari…