4 papers
Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
Jucheng Hu, Zhangquan Chen, Yulin Chen +9
Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical external signals (e.g., a cat's…
SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation
Sifei Li, Yang Li, Zizhou Wang +5
Cover songs constitute a vital aspect of musical culture, preserving the core melody of an original composition while reinterpreting it to infuse novel emotional depth and thematic…
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
Sifei Li, Mining Tan, Feier Shen +5
Multimodal learning has driven innovation across various industries, particularly in the field of music. By enabling more intuitive interaction experiences and enhancing immersion,…
VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
Sifei Li, Binxin Yang, Chunji Yin +4
Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achievi…