5 papers
CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space
Bowen Zhang, Junchuan Zhao, Ian McLoughlin +2
Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on s…
Peransformer: Improving Low-informed Expressive Performance Rendering with Score-aware Discriminator
Xian He, Wei Zeng, Ye Wang
Highly-informed Expressive Performance Rendering (EPR) systems transform music scores with rich musical annotations into human-like expressive performance MIDI files. While these s…
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
Wei Zeng, Junchuan Zhao, Ye Wang
Expressive performance rendering (EPR) and automatic piano transcription (APT) are fundamental yet inverse tasks in music information retrieval: EPR generates expressive performanc…
KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
Tianle Lyu, Junchuan Zhao, Ye Wang
Audio-driven facial animation has made significant progress in multimedia applications, with diffusion models showing strong potential for talking-face synthesis. However, most exi…
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
Junchuan Zhao, Wei Zeng, Tianle Lyu +1
Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete co…