13 papers
Towards Characterizing Scientific Image Utility and Upgradability
WenZhe Li, Qihang Yan, Liang Chen +6
Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content that introduces subtle but cons…
SIQA: Toward Reliable Scientific Image Quality Assessment
Wenzhe Li, Liang Chen, Junying Wang +6
Scientific images fundamentally differ from natural and AI-generated images in that they encode structured domain knowledge rather than merely depict visual scenes. Assessing their…
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
Xiaoxiao Wang, Chunxiao Li, Junying Wang +6
As comprehensive large model evaluation becomes prohibitively expensive, predicting model performance from limited observations has become essential. However, existing statistical…
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Ye Shen, Dun Pei, Yiqiu Guo +6
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…
QoNext: Towards Next-generation QoE for Foundation Models
Yijin Guo, Zicheng Zhang, Ye Shen +4
Existing evaluations of foundation models, including recent human-centric approaches, fail to capture what truly matters: user's experience during interaction. Current methods trea…
Improve MLLM Benchmark Efficiency through Interview
Farong Wen, Yijin Guo, Junying Wang +6
The rapid development of Multimodal Large Language Models (MLLM) has led to a wide range of MLLM applications, and a number of benchmark datasets have sprung up in order to assess…