5 citations · 10 across the 3 of their papers we have counts for
3 papers
eess.AS2023★ 5 cited
Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models
Guangzhi Sun, Wenyi Yu, Changli Tang +6
Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but…
cs.SD2023★ 1 cited
Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention Mechanism
Dichucheng Li, Mingjin Che, Wenwu Meng +4
Instrument playing technique (IPT) is a key element of musical presentation. However, most of the existing works for IPT detection only concern monophonic music signals, yet little…
cs.SD2022★ 4 cited
HPPNet: Modeling the Harmonic Structure and Pitch Invariance in Piano Transcription
Weixing Wei, Peilin Li, Yi Yu +1
While neural network models are making significant progress in piano transcription, they are becoming more resource-consuming due to requiring larger model size and more computing…