2 papers
cs.CR2025
VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
Ethan TS. Liu, Austin Wang, Spencer Mateega +2
Ensuring that large language models (LLMs) can effectively assess, detect, explain, and remediate software vulnerabilities is critical for building robust and secure software syste…
cs.LG2025
FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models
Spencer Mateega, Carlos Georgescu, Danny Tang
FinanceQA is a testing suite that evaluates LLMs' performance on complex numerical financial analysis tasks that mirror real-world investment work. Despite recent advances, current…