collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing

Xinlei Niu, Kin Wai Cheuk, Jing Zhang +8

Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing me…

cs.SD2025

Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech

Xinlei Niu, Jianbo Ma, Dylan Harper-Harris +3

The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on…

cs.SD2024

SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model

Xinlei Niu, Jing Zhang, Charles Patrick Martin

We present SoundMorpher, an open-world sound morphing method designed to generate perceptually uniform morphing trajectories. Traditional sound morphing techniques typically assume…

cs.SD2024

HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts

Xinlei Niu, Jing Zhang, Charles Patrick Martin

We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with cont…

cs.SD2024

SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation

Xinlei Niu, Jing Zhang, Christian Walder +1

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale…