2 papers
q-fin.ST2026
PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets
Avi Arora, Ritesh Malpani
Prediction markets offer a natural testbed for trading agents: contracts have binary payoffs, prices can be interpreted as probabilities, and realized performance depends criticall…
cs.SE2025
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Avi Arora, Jinu Jang, Roshanak Zilouchian Moghaddam
Modern Large Language Model (LLM) agents promise end to end assistance with real-world software tasks, yet existing benchmarks evaluate LLM agents almost exclusively in pre-baked e…