collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL

Yaxun Dai, Baolin Sun, Junying Wang +6

Tool-integrated Text-to-SQL parsing has emerged as a promising paradigm, framing SQL generation as a sequential decision-making process interleaved with tool execution. However, ex…

cs.CL2026

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Ye Shen, Dun Pei, Yiqiu Guo +6

Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…

cs.CL2025

QoNext: Towards Next-generation QoE for Foundation Models

Yijin Guo, Zicheng Zhang, Ye Shen +4

Existing evaluations of foundation models, including recent human-centric approaches, fail to capture what truly matters: user's experience during interaction. Current methods trea…

cs.CL2025

Improve MLLM Benchmark Efficiency through Interview

Farong Wen, Yijin Guo, Junying Wang +6

The rapid development of Multimodal Large Language Models (MLLM) has led to a wide range of MLLM applications, and a number of benchmark datasets have sprung up in order to assess…

cs.CL2025

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen +8

High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…

cs.CL2025

The Ever-Evolving Science Exam

Junying Wang, Zicheng Zhang, Yijin Guo +9

As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progr…