activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2025

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen +8

High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…

cs.CL2025

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Ye Shen, Junying Wang, Farong Wen +4

The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations…

cs.CL2025

QoNext: Towards Next-generation QoE for Foundation Models

Yijin Guo, Farong Wen, Ye Shen +5

Existing evaluations of foundation models predominantly focus on output correctness, treating interaction as a static exchange of information. However, such perspectives overlook t…

cs.CL2025

The Ever-Evolving Science Exam

Junying Wang, Zicheng Zhang, Yijin Guo +9

As foundation models grow rapidly in capability and deployment, evaluating their scientific understanding becomes increasingly critical. Existing science benchmarks have made progr…

cs.CL20252 cited

Research-Oriented Human-Centric Evaluation for Foundation Models

Yijin Guo, Kaiyuan Ji, Xiaorong Zhu +5

Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in h…

cs.CL2025

Improve MLLM Benchmark Efficiency through Interview

Farong Wen, Yijin Guo, Junying Wang +6

The rapid development of Multimodal Large Language Models (MLLM) has led to a wide range of MLLM applications, and a number of benchmark datasets have sprung up in order to assess…