5 papers · 1 filter
High-Fidelity Generative Audio Compression at 0.275kbps
Hao Ma, Ruihao Jing, Shansong Liu +4
High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditiona…
Towards Multimodal Query-Based Spatial Audio Source Extraction
Chenxin Yu, Hao Ma, Xu Li +4
Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leavin…
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Xueqing Li, Hao Ma, Zehan Li +8
Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
Hao Ma, Rujin Chen, Xiao-Lei Zhang +2
Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typ…
Language-Queried Target Sound Extraction Without Parallel Training Data
Hao Ma, Zhiyuan Peng, Xu Li +4
Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extens…