13 citations · 31 across the 8 of their papers we have counts for
13 papers
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Zhen Ye, Xinfa Zhu, Chi-Min Chan +17
Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and i…
Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders
Dichucheng Li, Yongyi Zang, Qiuqiang Kong
Automatic Music Transcription (AMT), aiming to get musical notes from raw audio, typically uses frame-level systems with piano-roll outputs or language model (LM)-based systems wit…
MusicScore: A Dataset for Music Score Modeling and Generation
Yuheng Lin, Zheqi Dai, Qiuqiang Kong
Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines,…
Joint Music and Language Attention Models for Zero-shot Music Tagging
Xingjian Du, Zhesong Yu, Jiaju Lin +2
Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be genera…
MERTech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model With Multi-Task Finetuning
Dichucheng Li, Yinghao Ma, Weixing Wei +6
Instrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled…
Transformer-based Autoencoder with ID Constraint for Unsupervised Anomalous Sound Detection
Jian Guan, Youde Liu, Qiuqiang Kong +4
Unsupervised anomalous sound detection (ASD) aims to detect unknown anomalous sounds of devices when only normal sound data is available. The autoencoder (AE) and self-supervised l…