5 papers
NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control
Yufan Wen, Zhaocheng Liu, YeGuo Hua +4
Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherenc…
A Mamba-Based Model for Automatic Chord Recognition
Chunyu Yuan, Johanna Devaney
In this work, we propose a new efficient solution, which is a Mamba-based model named BMACE (Bidirectional Mamba-based network, for Automatic Chord Estimation), which utilizes sele…
FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution
Junhao Zhuang, Shi Guo, Xin Cai +4
Diffusion models have recently advanced video restoration, but applying them to real-world video super-resolution (VSR) remains challenging due to high latency, prohibitive computa…
Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking
Zihan Su, Xuerui Qiu, Hongbin Xu +6
The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, inv…
TextureDiffusion: Target Prompt Disentangled Editing for Various Texture Transfer
Zihan Su, Junhao Zhuang, Chun Yuan
Recently, text-guided image editing has achieved significant success. However, existing methods can only apply simple textures like wood or gold when changing the texture of an obj…