activity
20242026
collaborators

8 papers

eess.AS2026

Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction

Zexu Pan, Xinyuan Qian, Shengkui Zhao +2

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human com…

eess.AS2025

Online Audio-Visual Autoregressive Speaker Extraction

Zexu Pan, Wupeng Wang, Shengkui Zhao +4

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less ex…

eess.AS2025

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

Zexu Pan, Shengkui Zhao, Tingting Wang +4

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. Ho…

cs.SD2025

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding

Dianwen Ng, Kun Zhou, Yi-Wen Chao +3

Achieving high-fidelity audio compression while preserving perceptual quality across diverse content remains a key challenge in Neural Audio Coding (NAC). We introduce MUFFIN, a fu…

cs.SD2025

InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation

Chong Zhang, Yukun Ma, Qian Chen +12

We introduce InspireMusic, a framework integrated super resolution and large language model for high-fidelity long-form music generation. A unified framework generates high-fidelit…

cs.SD2025

Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning

Shengkui Zhao, Zexu Pan, Kun Zhou +3

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on…