5 papers
FineVision: Open Data Is All You Need
Luis Wiedmann, Orr Zohar, Amir Mahla +6
The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduce FineVision, a meticulously col…
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
Shu Wu, Xiaotian Ye, Xinyu Mou +3
Multimodal knowledge editing (MKE) aims to correct the internal knowledge of large vision-language models after deployment, yet the behavioral patterns of post-edit models remain u…
GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing
Zihao Lin, Haibo Wang, Zhiyang Xu +7
Music-grounded mashup video creation is a challenging form of video non-linear editing, where a system must compose a coherent timeline from large collections of source videos whil…
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
Zhengyang Liang, Daoan Zhang, Huichi Zhou +8
While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilit…
Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps
Han Wang, Yilin Zhao, Dian Li +4
Humor is previously regarded as a gift exclusive to humans for the following reasons. Humor is a culturally nuanced aspect of human language, presenting challenges for its understa…