From the 1 of 9 linked papers with an AI index.
9 papers
Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
Zhiwei Lin, Jun Chen, Boshi Tang +5
Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous…
P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
Chong Jing, Junan Zhang, Jing Yang +3
MIDI-to-Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typicall…
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
Chong Jing, Junan Zhang, Jing Yang +3
The paper presents Anysynth, a diffusion‑transformer synthesizer that can render arbitrary target MIDI sequences with the timbre of an unseen instrument by directly conditioning on…
Schrödinger Bridge Mamba for One-Step Speech Enhancement
Jing Yang, Sirui Wang, Chao Wu +2
We present Schrödinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schrödinger Bridge (SB) training paradigm and the Mamba architecture.…
Multi-Metric Preference Alignment for Generative Speech Restoration
Junan Zhang, Xueyao Zhang, Jing Yang +3
Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptim…
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
Junan Zhang, Jing Yang, Zihao Fang +5
We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable o…