3 papers
cs.SE2026
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
cs.AI2025
What Limits Agentic Systems Efficiency?
Song Bian, Minghao Yan, Anand Jayarajan +2
Large Language Models (LLMs), such as OpenAI-o1 and DeepSeek-R1, have demonstrated strong reasoning capabilities. To further enhance LLM capabilities, recent agentic systems, such…
cs.LG2025
Scaling Inference-Efficient Language Models
Song Bian, Minghao Yan, Shivaram Venkataraman
Scaling laws are powerful tools to predict the performance of large language models. However, current scaling laws fall short of accounting for inference costs. In this work, we fi…