9 papers
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
Andong Deng, Dawei Du, Zhenfang Chen +7
Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. Whi…
D-Attn: Decomposed Attention for Large Vision-and-Language Models
Chia-Wen Kuo, Sijie Zhu, Fan Chen +2
Large vision-and-language models (LVLMs) have traditionally integrated visual and textual tokens by concatenating them into a single homogeneous input for large language models (LL…
Vidi: Large Multimodal Models for Video Understanding and Editing
Vidi Team, Celong Liu, Chia-Wen Kuo +20
Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support t…
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
Ming Li, Xin Gu, Fan Chen +4
Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signal…
Multi-Reward as Condition for Instruction-based Image Editing
Xin Gu, Ming Li, Libo Zhang +4
High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are c…
Where do Large Vision-Language Models Look at when Answering Questions?
Xiaoying Xing, Chia-Wen Kuo, Li Fuxin +6
Large Vision-Language Models (LVLMs) have shown promising performance in vision-language understanding and reasoning tasks. However, their visual understanding behaviors remain und…