5 citations · 6 across the 5 of their papers we have counts for
3 papers · 1 filter
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
Shansong Liu, Atin Sakkeer Hussain, Qilong Wu +2
Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored d…
MUGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Shansong Liu, Atin Sakkeer Hussain, Qilong Wu +2
The current landscape of research leveraging large language models (LLMs) is experiencing a surge. Many works harness the powerful reasoning capabilities of these models to compreh…
HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond
Shansong Liu, Xu Li, Dian Li +1
This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for down…