6 papers · 1 filter
VeRA: Verified Reasoning Data Augmentation at Scale
Zerui Cheng, Jiashuo Liu, Chunjie Wu +4
The main issue with most evaluation schemes today is their "static" nature: the same problems are reused repeatedly, allowing for memorization, format exploitation, and eventual sa…
FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains
Jiashuo Liu, Siyuan Chen, Zaiyuan Wang +38
Building upon FutureX, which established a live benchmark for general-purpose future prediction, this report introduces FutureX-Pro, including FutureX-Finance, FutureX-Retail, Futu…
When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
Zeshi Dai, Zimo Peng, Zerui Cheng +1
We present CAIA, a benchmark exposing a critical blind spot in AI evaluation: the inability of state-of-the-art models to operate in adversarial, high-stakes environments where mis…
Benchmarking is Broken -- Don't Let AI be its Own Judge
Zerui Cheng, Stella Wohnig, Ruchika Gupta +13
The meteoric rise of AI, with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need…
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
Jianzhu Yao, Kevin Wang, Ryan Hsieh +5
Reasoning and strategic behavior in social interactions is a hallmark of intelligence. This form of reasoning is significantly more sophisticated than isolated planning or reasonin…
OML: A Primitive for Reconciling Open Access with Owner Control in AI Model Distribution
Zerui Cheng, Edoardo Contente, Ben Finch +9
The current paradigm of AI model distribution presents a fundamental dichotomy: models are either closed and API-gated, sacrificing transparency and local execution, or openly dist…