2 papers
cs.LG2026
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
Philipp Mondorf, Samuel J. Bell, Jesse Dodge +1
As large language models (LLMs) are increasingly deployed to perform tasks with minimal human oversight, it is crucial that these models operate robustly. In particular, a model th…
cs.CL2026
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Yilun Zhao, Kaiyan Zhang, Tiansheng Hu +15
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific liter…