5 papers · 1 filter
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
Xinlei Niu, Kin Wai Cheuk, Jing Zhang +8
Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing me…
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
Xinlei Niu, Jianbo Ma, Dylan Harper-Harris +3
The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on…
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
Xinlei Niu, Jing Zhang, Charles Patrick Martin
We present SoundMorpher, an open-world sound morphing method designed to generate perceptually uniform morphing trajectories. Traditional sound morphing techniques typically assume…
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
Xinlei Niu, Jing Zhang, Charles Patrick Martin
We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with cont…
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
Xinlei Niu, Jing Zhang, Christian Walder +1
We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale…