5 papers · 1 filter
SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators
Yuheng Zhang, Yuanchun Wang, Fanjin Zhang +4
The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation…
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
Yelin Chen, Fanjin Zhang, Suping Sun +8
Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limite…
SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
Yuanchun Wang, Jifan Yu, Zijun Yao +13
Applying large language models (LLMs) for academic API usage shows promise in reducing researchers' academic information seeking efforts. However, current LLM API-using methods str…
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
Joy Lim Jia Yin, Daniel Zhang-Li, Jifan Yu +8
Evaluating the quality of slide-based multimedia instruction is challenging. Existing methods like manual assessment, reference-based metrics, and large language model evaluators f…
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
Shangqing Tu, Yuanchun Wang, Jifan Yu +6
Large language models have achieved remarkable success on general NLP tasks, but they may fall short for domain-specific problems. Recently, various Retrieval-Augmented Large Langu…