collaborators

6 papers

cs.CL2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Yilin Jiang, Xiaorong Zhu, Fei Tan +9

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…

cs.AI2026

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

Ziheng Jia, Jiaying Qian, Zicheng Zhang +3

AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation model…

cs.CV2025

GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs

Xiaorong Zhu, Ziheng Jia, Jiarui Wang +6

The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assess…

cs.CV2025

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models

Jiarui Wang, Huiyu Duan, Juntong Wang +8

With the rapid advancement of generative models, the realism of AI-generated images has significantly improved, posing critical challenges for verifying digital content authenticit…

cs.CL2025

Research-Oriented Human-Centric Evaluation for Foundation Models

Yijin Guo, Kaiyuan Ji, Xiaorong Zhu +5

Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in h…

cs.CV2025

Scaling-up Perceptual Video Quality Assessment

Ziheng Jia, Zicheng Zhang, Zeyu Zhang +12

The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptua…