4 papers
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
Tu Trinh, Mohamed Elfeki, Guangze Luo +9
Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete or ambiguous. The bottleneck is not raw capability, but judgm…
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie +14
Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While ex…
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi +13
Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities,…
Remote Labor Index: Measuring AI Automation of Remote Work
Mantas Mazeika, Alice Gatti, Cristina Menghini +44
AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation. To mea…