activity
20242026
collaborators

11 papers

cs.SD2026

Scalable Neural Vocoder from Range-Null Space Decomposition

Andong Li, Tong Lei, Zhihang Sun +4

Although deep neural networks have facilitated significant progress of neural vocoders in recent years, they usually suffer from intrinsic challenges like opaque modeling, inflexib…

eess.AS2026

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

Yifan Liang, Andong Li, Kang Yang +5

Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as me…

cs.SD2026

GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks

Lingling Dai, Andong Li, Cheng Chi +3

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR…

cs.SD2025

BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective

Andong Li, Tong Lei, Rilin Chen +5

This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare…

cs.SD2025

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

Kang Yang, Yifan Liang, Fangkun Liu +2

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To…

eess.AS2025

Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders

Lingling Dai, Andong Li, Tong Lei +3

Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively…