2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.AI2026
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
Chris Ge, Daria Kryvosheieva, Daniel Fried +2
As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challe…
cs.LG2025
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents
Kaivalya Hariharan, Uzay Girit, Atticus Wang +1
Benchmarks for large language models (LLMs) have predominantly assessed short-horizon, localized reasoning. Existing long-horizon suites (e.g. SWE-bench) rely on manually curated i…
cs.CL2024★ 2 cited
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
Rudolf Laine, Bilal Chughtai, Jan Betley +6
AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model". This raises questions. Do such models know that they are LLMs and reliably a…