3 papers
cs.LG2026
AllocBench: Measuring Online Tool Allocation Capability in LLM Agents
Daniel Wang, Andrew Xu
Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a sma…
cs.AI2026
Lomekwi: Resource-Bounded Tool Discovery in LLM Agents
Roshan Klein-Seetharaman, Daniel Wang, Andrew Xu
Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distinguish tool use from tool discovery and de…
cs.SE2026
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…