3 papers
cs.CL2026
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Zhipeng Xu, Jiahao Lu, Yining Zheng +2
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evide…
cs.CL2026
What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?
Chuanyuan Tan, Junjie Yu, Yuxin Wang +3
Reliable handling of unanswerable questions (UAQs) is critical for trustworthy LLM-based agents. Although memory is widely used in agent systems, its role in reliable UAQ handling…
cs.CL2026
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
Bowen Li, Haochen Ma, Yuxin Wang +5
The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primaril…