collaborators

6 papers

cs.SE2026

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Jian Zhu, Yuzheng Zhang, Zeyao Ma +11

Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations s…

cs.SE2026

IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks

Spencer Mateega, Jeff Yang, Tiana Costello +3

IDE-Bench is a comprehensive framework for evaluating AI IDE agents on real-world software engineering tasks through an IDE-native tool interface. We present a Dockerized test harn…

cs.CL2026

Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics

Abhay Srivastava, Sam Jung, Spencer Mateega

We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters fro…

cs.CL2025

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

Sam Jung, Agustin Garcinuno, Spencer Mateega

AI text-to-app tools promise high quality applications and websites in minutes, yet no public benchmark rigorously verifies those claims. We introduce UI-Bench, the first large-sca…

cs.CR2025

VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation

Ethan TS. Liu, Austin Wang, Spencer Mateega +2

Ensuring that large language models (LLMs) can effectively assess, detect, explain, and remediate software vulnerabilities is critical for building robust and secure software syste…

cs.LG2025

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Spencer Mateega, Carlos Georgescu, Danny Tang

FinanceQA is a testing suite that evaluates LLMs' performance on complex numerical financial analysis tasks that mirror real-world investment work. Despite recent advances, current…