5 papers
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
Shaoxuan Li, Zhixuan Zhao, Hanze Deng +9
We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is su…
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
Tian Lan, Yang-Hao Zhou, Zi-Ao Ma +8
Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these gener…
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu +5
The rapid progress in diffusion-based text-to-image (T2I) generation has created an urgent need for interpretable automatic evaluation methods that can assess the quality of genera…
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu +6
We present a systematic investigation of Multi-modal Retrieval Augmented Multi-modal Generation (MRAG), a novel task that enables foundation models to process multi-modal web c…
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
Rong-Cheng Tu, Zi-Ao Ma, Tian Lan +3
Driven by the remarkable progress in diffusion models, text-to-image generation has made significant strides, creating a pressing demand for automatic quality evaluation of generat…