5 papers
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Ye Shen, Dun Pei, Yiqiu Guo +6
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
Qi Jia, Ye Shen, Xiujie Song +5
Evaluating LLMs' instruction-following ability in multi-topic dialogues is essential yet challenging. Existing benchmarks are limited to a fixed number of turns, susceptible to sat…
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
Junying Wang, Zicheng Zhang, Ye Shen +8
High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
Ye Shen, Junying Wang, Farong Wen +4
The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations…
The Ever-Evolving Science Exam
Junying Wang, Zicheng Zhang, Yijin Guo +9
As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progr…