4 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
What Limits Agentic Systems Efficiency?
Song Bian, Minghao Yan, Anand Jayarajan +2
Large Language Models (LLMs), such as OpenAI-o1 and DeepSeek-R1, have demonstrated strong reasoning capabilities. To further enhance LLM capabilities, recent agentic systems, such…
Tesserae: Scalable Placement Policies for Deep Learning Workloads
Song Bian, Saurabh Agarwal, Md. Tareq Mahmood +1
Training deep learning (DL) models has become a dominant workload in data-centers and improving resource utilization is a key goal of DL cluster schedulers. In order to do this, sc…
Scaling Inference-Efficient Language Models
Song Bian, Minghao Yan, Shivaram Venkataraman
Scaling laws are powerful tools to predict the performance of large language models. However, current scaling laws fall short of accounting for inference costs. In this work, we fi…