2 citations · 4 across the 5 of their papers we have counts for
5 papers
HumTrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond
Shansong Liu, Xu Li, Dian Li +1
This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for down…
Unified Pretraining Target Based Video-music Retrieval With Music Rhythm And Video Optical Flow Information
Tianjun Mao, Shansong Liu, Yunxuan Zhang +2
Background music (BGM) can enhance the video's emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrie…
Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun +1
Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we p…
A Hierarchical Speaker Representation Framework for One-shot Singing Voice Conversion
Xu Li, Shansong Liu, Ying Shan
Typically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speake…
Neural Architecture Search For LF-MMI Trained Time Delay Neural Networks
Shoukang Hu, Xurong Xie, Mingyu Cui +6
State-of-the-art automatic speech recognition (ASR) system development is data and computation intensive. The optimal design of deep neural networks (DNNs) for these systems often…