From the 1 of 9 linked papers with an AI index.
9 papers
P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
Chong Jing, Junan Zhang, Jing Yang +3
MIDI-to-Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typicall…
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
Chong Jing, Junan Zhang, Jing Yang +3
The paper presents Anysynth, a diffusion‑transformer synthesizer that can render arbitrary target MIDI sequences with the timbre of an unseen instrument by directly conditioning on…
Aliasing-Free Neural Audio Synthesis
Yicheng Gu, Junan Zhang, Chaoren Wang +3
In neural audio synthesis, neural vocoders and codecs are models that reconstruct waveforms from acoustic and latent representations, which are essential to the resulting audio qua…
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
Lester Phillip Violeta, Xueyao Zhang, Jiatong Shi +4
We present a thorough analysis of the findings of the latest iteration of the Singing Voice Conversion Challenge, a scientific event aiming to compare and understand different voic…
EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction
Chong Jing, Zitong Lan, Junan Zhang +1
Predicting spatially varying Room Impulse Response (RIR) from sparse observations is a critical but highly challenging inverse problem for immersive spatial audio rendering. In thi…
The CCF AATC 2025 Speech Restoration Challenge: A Retrospective
Junan Zhang, Mengyao Zhu, Xin Xu +3
Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, i…