2 papers
cs.CL2026
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
Yuchen Lu, Run Yang, Yichen Zhang +6
Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do not reflect the layered, proof-dri…
cs.CL2026
SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation
Beichen Guo, Zhiyuan Wen, Jia Gu +6
Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialize…