2 citations · 3 across the 4 of their papers we have counts for
5 papers · 1 filter
GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models
Zuyao You, Zhesong Yu, Mingyu Liu +3
In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamli…
ByteComposer: a Human-like Melody Composition Method based on Language Model Agent
Xia Liang, Xingjian Du, Jiaju Lin +3
Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody compos…
Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding
Chao Wang, Zhonghao Li, Benlai Tang +4
Recently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in…
PPG-based singing voice conversion with adversarial representation learning
Zhonghao Li, Benlai Tang, Xiang Yin +4
Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion work…
High-resolution Piano Transcription with Pedals by Regressing Onset and Offset Times
Qiuqiang Kong, Bochen Li, Xuchen Song +2
Automatic music transcription (AMT) is the task of transcribing audio recordings into symbolic representations. Recently, neural network-based methods have been applied to AMT, and…