26 papers
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
Qi Jia, Haodong Zhao, Dun Pei +7
Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities. However, existing evaluation p…
GeoR-Bench: Evaluating Geoscience Visual Reasoning
Yushuo Zheng, Zicheng Zhang, Huiyu Duan +7
Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, cl…
SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond
Xiangyang Zhu, Yuan Tian, Qi Jia +14
The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchm…
A: Towards Advertising Aesthetic Assessment
Kaiyuan Ji, Yixuan Gao, Lu Sun +7
Advertising images significantly impact commercial conversion rates and brand equity, yet current evaluation methods rely on subjective judgments, lacking scalability, standardized…
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
Xiaoxiao Wang, Chunxiao Li, Junying Wang +6
As comprehensive large model evaluation becomes prohibitively expensive, predicting model performance from limited observations has become essential. However, existing statistical…
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
Zijian Chen, Wenjun Zhang, Guangtao Zhai
The potential data contamination issue in contemporary large language models (LLMs) benchmarks presents a fundamental challenge to establishing trustworthy evaluation frameworks. M…