collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

High-Fidelity Generative Audio Compression at 0.275kbps

Hao Ma, Ruihao Jing, Shansong Liu +4

High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditiona…

eess.AS2025

Towards Multimodal Query-Based Spatial Audio Source Extraction

Chenxin Yu, Hao Ma, Xu Li +4

Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leavin…

eess.AS2025

Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization

Xueqing Li, Hao Ma, Zehan Li +8

Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…

eess.AS2025

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Hao Ma, Rujin Chen, Xiao-Lei Zhang +2

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typ…

eess.AS2024

Language-Queried Target Sound Extraction Without Parallel Training Data

Hao Ma, Zhiyuan Peng, Xu Li +4

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extens…