16 papers
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
Bingxuan Li, Yiming Cui, Yicheng He +4
Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-…
VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
Yexiang Liu, Wen Zhong, Sijie Zhu +6
The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is high…
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
Andong Deng, Dawei Du, Zhenfang Chen +7
Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. Whi…
Referring Layer Decomposition
Fangyi Chen, Yaojie Shen, Lu Xu +4
Precise, object-aware control over visual content is essential for advanced image editing and compositional generation. Yet, most existing approaches operate on entire images holis…
Vidi2.5: Large Multimodal Models for Video Understanding and Creation
Vidi Team, Chia-Wen Kuo, Chuang Huang +31
Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to…
Structured Context Learning for Generic Event Boundary Detection
Xin Gu, Congcong Li, Xinyao Wang +5
Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, ca…