From the 3 of 9 linked papers with an AI index.
9 papers
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Tengfei Liu, Yang Shi, Yuran Wang +16
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Yuqing Wen, Yukai Huang, Qianqian Xie +6
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated…
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Yuqi Tang, Tengfei Liu, Yizheng Lai +18
The paper introduces KeyFrame-Compass, a benchmark and evaluation framework for assessing how well video generation models follow supplied keyframes while preserving overall video…
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Xiaohan Zhang, Yuqing Wen, Junlin Chen +9
The paper introduces MultiRef-Compass, a benchmark for evaluating models that generate synchronized audio‑video content conditioned on multiple references and textual instructions,…
CoVEBench: Can Video Editing Models Handle Complex Instructions?
Jiangtao Wu, Jiaming Wang, Yiwen He +7
While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly compositional. A single prompt…
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
Zhe Cao, Tao Wang, Jiaming Wang +10
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented,…