works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.SD2026

P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing

Chong Jing, Junan Zhang, Jing Yang +3

MIDI-to-Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typicall…

cs.SD2026

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

Chong Jing, Junan Zhang, Jing Yang +3

The paper presents Anysynth, a diffusion‑transformer synthesizer that can render arbitrary target MIDI sequences with the timbre of an unseen instrument by directly conditioning on…

eess.AS2026

A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic

Jing Yang, Shuqing Zhang, Yongyi Deng +5

Accurate phoneme recognition is pivotal for mispronunciation detection and diagnosis (MDD) in modern standard Arabic (MSA), yet remains constrained by data scarcity and the synthet…

cs.SD2026

Schrödinger Bridge Mamba for One-Step Speech Enhancement

Jing Yang, Sirui Wang, Chao Wu +2

We present Schrödinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schrödinger Bridge (SB) training paradigm and the Mamba architecture.…

cs.SD2025

Multi-Metric Preference Alignment for Generative Speech Restoration

Junan Zhang, Xueyao Zhang, Jing Yang +3

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptim…

cs.SD2025

AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement

Junan Zhang, Jing Yang, Zihao Fang +5

We introduce AnyEnhance, a unified generative model for voice enhancement that processes both speech and singing voices. Based on a masked generative model, AnyEnhance is capable o…