activity
20232026
collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2026

P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing

Chong Jing, Junan Zhang, Jing Yang +3

MIDI-to-Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typicall…

cs.SD2026

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

Chong Jing, Junan Zhang, Jing Yang +3

Zero-shot instrument cloning aims to render an arbitrary [Target MIDI] sequence with the acoustic identity of an unseen instrument given only a short [Reference Audio, Reference MI…

cs.SD2025

BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model

Ganghui Ru, Jieying Wang, Jiahao Zhao +5

Beat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits…

cs.SD2025

HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking

Ganghui Ru, Jieying Wang, Jiahao Zhao +5

Fine-tuning pre-trained foundation models has made significant progress in music information retrieval. However, applying these models to beat tracking tasks remains unexplored as…

cs.SD2025

Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection

Weixing Wei, Jiahao Zhao, Yulun Wu +1

This paper describes a streaming audio-to-MIDI piano transcription approach that aims to sequentially translate a music signal into a sequence of note onset and offset events. The…

cs.SD2023

Ms-senet: Enhancing Speech Emotion Recognition Through Multi-scale Feature Fusion With Squeeze-and-excitation Blocks

Mengbo Li, Yuanzhong Zheng, Dichucheng Li +3

Speech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. Spatiotemporal features play a crucial role in SER, yet current research lack…