most citedCWS-PResUNet: Music Source Separation with Channel-wise Subband Phase-aware ResUNet

13 citations · 31 across the 8 of their papers we have counts for

collaborators

13 papers

eess.AS2025

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Zhen Ye, Xinfa Zhu, Chi-Min Chan +17

Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and i…

cs.SD2025

Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders

Dichucheng Li, Yongyi Zang, Qiuqiang Kong

Automatic Music Transcription (AMT), aiming to get musical notes from raw audio, typically uses frame-level systems with piano-roll outputs or language model (LM)-based systems wit…

cs.MM20241 cited

MusicScore: A Dataset for Music Score Modeling and Generation

Yuheng Lin, Zheqi Dai, Qiuqiang Kong

Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines,…

cs.SD2023

Joint Music and Language Attention Models for Zero-shot Music Tagging

Xingjian Du, Zhesong Yu, Jiaju Lin +2

Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be genera…

cs.SD2023

MERTech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model With Multi-Task Finetuning

Dichucheng Li, Yinghao Ma, Weixing Wei +6

Instrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled…

cs.SD2023

Transformer-based Autoencoder with ID Constraint for Unsupervised Anomalous Sound Detection

Jian Guan, Youde Liu, Qiuqiang Kong +4

Unsupervised anomalous sound detection (ASD) aims to detect unknown anomalous sounds of devices when only normal sound data is available. The autoencoder (AE) and self-supervised l…