3 papers
cs.CV2026
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering
Mouxiao Huang, Qiangyu Yan, Borui Jiang +1
Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and L…
cs.CV2025
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
Yiman Zhang, Ziheng Luo, Qiangyu Yan +4
In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with exis…
cs.CV2025
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
Zhenliang Ni, Qiangyu Yan, Mouxiao Huang +5
The rapid advancement of video generation models has made it increasingly challenging to distinguish AI-generated videos from real ones. This issue underscores the urgent need for…