3 papers
eess.AS2026
SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis
Yifan Liang, Andong Li, Kang Yang +5
Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as me…
cs.SD2026
GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks
Lingling Dai, Andong Li, Cheng Chi +3
In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR…
eess.AS2025
Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
Lingling Dai, Andong Li, Tong Lei +3
Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively…