most citedMusic Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

2 citations · 4 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2023

HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond

Shansong Liu, Xu Li, Dian Li +1

This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for down…

cs.MM20231 cited

Unified Pretraining Target Based Video-music Retrieval With Music Rhythm And Video Optical Flow Information

Tianjun Mao, Shansong Liu, Yunxuan Zhang +2

Background music (BGM) can enhance the video's emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrie…

cs.SD20232 cited

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun +1

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we p…

eess.AS2022

A Hierarchical Speaker Representation Framework for One-shot Singing Voice Conversion

Xu Li, Shansong Liu, Ying Shan

Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speake…

eess.AS20221 cited

Neural Architecture Search For LF-MMI Trained Time Delay Neural Networks

Shoukang Hu, Xurong Xie, Mingyu Cui +6

State-of-the-art automatic speech recognition (ASR) system development is data and computation intensive. The optimal design of deep neural networks (DNNs) for these systems often…