2 papers
cs.CL2026
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Xinming Tu, Tianze Wang, Yingzhou +4
As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit ass…
cs.AI2025
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Ming Yin, Yuanhao Qu, Ling Yang +2
We investigate how to teach large language models (LLMs) to perform scientific reasoning by leveraging expert discussions as a learning signal. Focusing on the genomics domain, we…