collaborators

5 papers

cs.CL2026

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Ye Shen, Dun Pei, Yiqiu Guo +6

Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…

cs.CL2025

One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework

Qi Jia, Ye Shen, Xiujie Song +5

Evaluating LLMs' instruction-following ability in multi-topic dialogues is essential yet challenging. Existing benchmarks are limited to a fixed number of turns, susceptible to sat…

cs.CL2025

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen +8

High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…

cs.CL2025

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Ye Shen, Junying Wang, Farong Wen +4

The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations…

cs.CL2025

The Ever-Evolving Science Exam

Junying Wang, Zicheng Zhang, Yijin Guo +9

As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progr…