4 papers
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
Xueqing Li, Hao Ma, Zehan Li +8
Self-supervised learning (SSL) has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving eff…
FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
Lei Zhao, Rujin Chen, Chi Zhang +2
Recently, with the advancement of AIGC, deep learning-based video-to-audio (V2A) technology has garnered significant attention. However, existing research mostly focuses on mono au…
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Lei Zhao, Linfeng Feng, Dongxu Ge +5
With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited explorati…
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
Hao Ma, Rujin Chen, Xiao-Lei Zhang +2
Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typ…