2 papers
cs.CL2026
Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics
Abhay Srivastava, Sam Jung, Spencer Mateega
We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters fro…
cs.CL2025
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
Sam Jung, Agustin Garcinuno, Spencer Mateega
AI text-to-app tools promise high quality applications and websites in minutes, yet no public benchmark rigorously verifies those claims. We introduce UI-Bench, the first large-sca…