18 papers
Emo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment
Lancheng Gao, Ziheng Jia, Shengyan Li +4
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benc…
Visual Distortion Detection in UGC Images Using Large Multimodal Models
Ziheng Jia, Yingji Liang, Jiaying Qian +1
The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal…
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
Ziheng Jia, Jiaying Qian, Zicheng Zhang +3
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation model…
DyCoRM: Dynamic Criterion-Aware Reward Modeling for Text-to-Image Generation
Jiaying Qian, Ziheng Jia, Qian Zhang +5
With the continued advancement of text-to-image (T2I) generation, producing high-quality images is becoming increasingly attainable; consequently, user demands are shifting toward…
GeoR-Bench: Evaluating Geoscience Visual Reasoning
Yushuo Zheng, Zicheng Zhang, Huiyu Duan +7
Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, cl…
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
Ziheng Jia, Linhan Cao, Jinliang Han +6
Developing a robust visual quality assessment (VQualA) large multi-modal model (LMM) requires achieving versatility, powerfulness, and transferability. However, existing VQualA LMM…