2 papers
cs.AI2026
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
Chris Ge, Daria Kryvosheieva, Daniel Fried +2
As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challe…
cs.LG2025
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents
Kaivalya Hariharan, Uzay Girit, Atticus Wang +1
Benchmarks for large language models (LLMs) have predominantly assessed short-horizon, localized reasoning. Existing long-horizon suites (e.g. SWE-bench) rely on manually curated i…