Showing eess.ASShow all
2 papers · 1 filter
eess.AS2026
SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis
Yifan Liang, Andong Li, Kang Yang +5
Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as me…
eess.AS2025
Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
Lingling Dai, Andong Li, Tong Lei +3
Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively…