1 citations · 1 across the 5 of their papers we have counts for
8 papers · 1 filter
MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
Shuyu Li, Kejun Zhang, Jiahe Lei +5
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and…
Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation
Shulei Ji, Zihao Wang, Jiaxing Yu +4
Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm mode…
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
Shuyu Li, Shulei Ji, Zihao Wang +3
Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applica…
MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
Haoxuan Liu, Zihao Wang, Haorong Hong +5
This paper introduces MetaBGM, a groundbreaking framework for generating background music that adapts to dynamic scenes and real-time user interactions. We define multi-scene as va…
MuDiT & MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation
Zihao Wang, Haoxuan Liu, Jiaxing Yu +3
Amid the rising intersection of generative AI and human artistic processes, this study probes the critical yet less-explored terrain of alignment in human-centric automatic song co…
SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement
Zihao Wang, Le Ma, Yongsheng Feng +3
Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can h…