6 papers
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Jian Zhu, Yuzheng Zhang, Zeyao Ma +11
Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations s…
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
Spencer Mateega, Jeff Yang, Tiana Costello +3
IDE-Bench is a comprehensive framework for evaluating AI IDE agents on real-world software engineering tasks through an IDE-native tool interface. We present a Dockerized test harn…
Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics
Abhay Srivastava, Sam Jung, Spencer Mateega
We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters fro…
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
Sam Jung, Agustin Garcinuno, Spencer Mateega
AI text-to-app tools promise high quality applications and websites in minutes, yet no public benchmark rigorously verifies those claims. We introduce UI-Bench, the first large-sca…
VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
Ethan TS. Liu, Austin Wang, Spencer Mateega +2
Ensuring that large language models (LLMs) can effectively assess, detect, explain, and remediate software vulnerabilities is critical for building robust and secure software syste…
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
Spencer Mateega, Carlos Georgescu, Danny Tang
FinanceQA is a testing suite that evaluates LLMs' performance on complex numerical financial analysis tasks that mirror real-world investment work. Despite recent advances, current…