From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Zhongwei Ren, Yunchao Wei, Yao Zhao +5
The paper presents UniVR, a system that learns complex visual reasoning, fine-grained physical dynamics, and long-term planning directly from raw video demonstrations using a novel…
cs.CV2024
VCoME: Verbal Video Composition with Multimodal Editing Effects
Weibo Gong, Xiaojie Jin, Xin Li +2
Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to…
cs.CV2024
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
Xiaojie Jin, Bowen Zhang, Weibo Gong +6
State-of-the-art video-text retrieval (VTR) methods typically involve fully fine-tuning a pre-trained model (e.g. CLIP) on specific datasets. However, this can result in significan…