collaborators

16 papers

cs.CV2026

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

Bingxuan Li, Yiming Cui, Yicheng He +4

Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-…

cs.AI2026

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

Yexiang Liu, Wen Zhong, Sijie Zhu +6

The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is high…

cs.CV2026

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Andong Deng, Dawei Du, Zhenfang Chen +7

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. Whi…

cs.CV2026

Referring Layer Decomposition

Fangyi Chen, Yaojie Shen, Lu Xu +4

Precise, object-aware control over visual content is essential for advanced image editing and compositional generation. Yet, most existing approaches operate on entire images holis…

cs.CV2026

Vidi2.5: Large Multimodal Models for Video Understanding and Creation

Vidi Team, Chia-Wen Kuo, Chuang Huang +31

Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to…

cs.CV2025

Structured Context Learning for Generic Event Boundary Detection

Xin Gu, Congcong Li, Xinyao Wang +5

Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, ca…