10 citations · 31 across the 9 of their papers we have counts for
12 papers
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
Haonan Chen, Jordan B. L. Smith, Janne Spijkervet +5
Progress in the task of symbolic music generation may be lagging behind other tasks like audio and text generation, in part because of the scarcity of symbolic training data. In th…
Graph Contrastive Learning with Implicit Augmentations
Huidong Liang, Xingjian Du, Bilei Zhu +3
Existing graph contrastive learning methods rely on augmentation techniques based on random perturbations (e.g., randomly adding or dropping edges and nodes). Nevertheless, alterin…
HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
Ke Chen, Xingjian Du, Bilei Zhu +3
Audio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted…
Attention-based cross-modal fusion for audio-visual voice activity detection in musical video streams
Yuanbo Hou, Zhesong Yu, Xia Liang +4
Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse mu…
Speech enhancement with weakly labelled data from AudioSet
Qiuqiang Kong, Haohe Liu, Xingjian Du +3
Speech enhancement is a task to improve the intelligibility and perceptual quality of degraded speech signal. Recently, neural networks based methods have been applied to speech en…
CatNet: music source separation system with mix-audio augmentation
Xuchen Song, Qiuqiang Kong, Xingjian Du +1
Music source separation (MSS) is the task of separating a music piece into individual sources, such as vocals and accompaniment. Recently, neural network based methods have been ap…