collaborators

6 papers

cs.CL2025

MSRS: Evaluating Multi-Source Retrieval-Augmented Generation

Rohan Phanse, Yijie Zhou, Kejian Shi +4

Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or fa…

cs.CL2025

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research

Yilun Zhao, Weiyuan Chen, Zhijian Xu +5

We introduce AbGen, the first benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research. AbGen consists of 1,500 expert-annotate…

cs.CL2025

Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

Mingqi Gao, Yixin Liu, Xinyu Hu +3

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consumi…

cs.CV2025

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Yilun Zhao, Lujing Xie, Haowei Zhang +16

We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions…

cs.CL2024

M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models

Chuhan Li, Ziyao Shangguan, Yilun Zhao +3

Existing benchmarks for evaluating foundation models mainly focus on single-document, text-only tasks. However, they often fail to fully capture the complexity of research workflow…

cs.CL2024

ReIFE: Re-evaluating Instruction-Following Evaluation

Yixin Liu, Kejian Shi, Alexander R. Fabbri +5

The automatic evaluation of instruction following typically involves using large language models (LLMs) to assess response quality. However, there is a lack of comprehensive evalua…