5 papers
MotionBeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding
Xuanchen Wang, Heng Wang, Weidong Cai
Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect thi…
MusicWeaver: Programmable Long-Form Music Generation with Provably Local Editing
Xuanchen Wang, Heng Wang, Weidong Cai
Music generation systems produce increasingly realistic audio, yet they expose no interface between a creator's structural intent and the rendered sound. Form can only be steered t…
Collapse of Patches: Ranking Image Patches for Efficient Visual Modeling
Wei Guo, Shunqi Mao, Zhuonan Liang +3
Observing certain patches in an image reduces the uncertainty of others. Their realization lowers the distribution entropy of each remaining patch feature, analogous to collapsing…
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
Wei Guo, Heng Wang, Jianbo Ma +1
Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or…
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
Xuanchen Wang, Heng Wang, Weidong Cai
Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches o…