3 papers
cs.SE2026
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
cs.CL2025
LangProBe: a Language Programs Benchmark
Shangyin Tan, Lakshya A Agrawal, Arnav Singhvi +6
Composing language models (LMs) into multi-step language programs and automatically optimizing their modular prompts is now a mainstream paradigm for building AI systems, but the t…
cs.LG2025
S*: Test Time Scaling for Code Generation
Dacheng Li, Shiyi Cao, Chengkun Cao +6
Increasing test-time compute for LLMs shows promise across domains but remains underexplored in code generation, despite extensive study in math. In this paper, we propose S*, the…