9 papers · 1 filter
Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL
Yaxun Dai, Baolin Sun, Junying Wang +6
Tool-integrated Text-to-SQL parsing has emerged as a promising paradigm, framing SQL generation as a sequential decision-making process interleaved with tool execution. However, ex…
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Ye Shen, Dun Pei, Yiqiu Guo +6
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…
QoNext: Towards Next-generation QoE for Foundation Models
Yijin Guo, Zicheng Zhang, Ye Shen +4
Existing evaluations of foundation models, including recent human-centric approaches, fail to capture what truly matters: user's experience during interaction. Current methods trea…
Improve MLLM Benchmark Efficiency through Interview
Farong Wen, Yijin Guo, Junying Wang +6
The rapid development of Multimodal Large Language Models (MLLM) has led to a wide range of MLLM applications, and a number of benchmark datasets have sprung up in order to assess…
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
Junying Wang, Zicheng Zhang, Ye Shen +8
High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…
The Ever-Evolving Science Exam
Junying Wang, Zicheng Zhang, Yijin Guo +9
As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progr…