6 papers
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Yilin Jiang, Xiaorong Zhu, Fei Tan +9
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
Ziheng Jia, Jiaying Qian, Zicheng Zhang +3
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation model…
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
Xiaorong Zhu, Ziheng Jia, Jiarui Wang +6
The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assess…
DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
Jiarui Wang, Huiyu Duan, Juntong Wang +8
With the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticit…
Research-Oriented Human-Centric Evaluation for Foundation Models
Yijin Guo, Kaiyuan Ji, Xiaorong Zhu +5
Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in h…
Scaling-up Perceptual Video Quality Assessment
Ziheng Jia, Zicheng Zhang, Zeyu Zhang +12
The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptua…