36 papers
Towards Characterizing Scientific Image Utility and Upgradability
WenZhe Li, Qihang Yan, Liang Chen +6
Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content that introduces subtle but cons…
GeoR-Bench: Evaluating Geoscience Visual Reasoning
Yushuo Zheng, Zicheng Zhang, Huiyu Duan +7
Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, cl…
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
Shengyu Guo, Tongrui Ye, Jianbo Zhang +3
Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence.…
SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond
Xiangyang Zhu, Yuan Tian, Qi Jia +14
The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchm…
Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content
Shushi Wang, Zicheng Zhang, Chunyi Li +7
Quality assessment of AI-generated content is crucial for evaluating model capability and guiding model optimization. However, most existing quality assessment datasets and models…
SIQA: Toward Reliable Scientific Image Quality Assessment
Wenzhe Li, Liang Chen, Junying Wang +6
Scientific images fundamentally differ from natural and AI-generated images in that they encode structured domain knowledge rather than merely depict visual scenes. Assessing their…