18 citations · 18 across the 2 of their papers we have counts for
2 papers
cs.SE2026
Evaluating Agentic Code Repair Capabilities in Distributed Systems
Yibo Yan, Huijuan Wang, Junzhou He +4
LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging,…
cs.LG2026★ 18 cited
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…