4 papers
CutClaw: Agentic Hours-Long Video Editing via Music Synchronization
Shifang Zhao, Yihan Hu, Ying Shan +2
Editing the video content with audio alignment forms a digital human-made art in current social media. However, the time-consuming and repetitive nature of manual video editing has…
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
Siyuan Hou, Shansong Liu, Ruibin Yuan +4
Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram repr…
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
Shansong Liu, Atin Sakkeer Hussain, Qilong Wu +2
Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored d…
MUGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Shansong Liu, Atin Sakkeer Hussain, Qilong Wu +2
The current landscape of research leveraging large language models (LLMs) is experiencing a surge. Many works harness the powerful reasoning capabilities of these models to compreh…