collaborators

7 papers

cs.CL2025

MSRS: Evaluating Multi-Source Retrieval-Augmented Generation

Rohan Phanse, Yijie Zhou, Kejian Shi +4

Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or fa…

cs.CL2025

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research

Yilun Zhao, Weiyuan Chen, Zhijian Xu +5

We introduce AbGen, the first benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research. AbGen consists of 1,500 expert-annotate…

cs.AI2025

PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving

Kaiyue Feng, Yilun Zhao, Yixin Liu +4

We introduce PHYSICS, a comprehensive benchmark for university-level physics problem solving. It contains 1297 expert-annotated problems covering six core areas: classical mechanic…

cs.CL2025

Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

Mingqi Gao, Yixin Liu, Xinyu Hu +3

Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consumi…

cs.CV2025

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Yilun Zhao, Lujing Xie, Haowei Zhang +16

We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions…

cs.CL2024

M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models

Chuhan Li, Ziyao Shangguan, Yilun Zhao +3

Existing benchmarks for evaluating foundation models mainly focus on single-document, text-only tasks. However, they often fail to fully capture the complexity of research workflow…