collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Xinyu Liu, Shihao Li, Weihong Lin +10

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordi…

cs.CV2026

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

Xinlong Chen, Jiafu Tang, Yue Ding +12

Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can ob…

cs.CV2026

CoVEBench: Can Video Editing Models Handle Complex Instructions?

Jiangtao Wu, Jiaming Wang, Yiwen He +7

While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional. A single prompt…

cs.CV2026

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Jiahao Wang, An Ping, Yanghai Wang +13

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…

cs.CV2025

ViDiC: Video Difference Captioning

Jiangtao Wu, Shihao Li, Zhaozhou Bian +7

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored…

cs.CV2025

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

Tianhao Peng, Haochen Wang, Yuanxing Zhang +13

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understa…