3 papers
cs.AI2026
BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents
Alex Wang, Georg Meinhardt, Jacob Katz +4
Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were us…
cs.AI2025
The Leaderboard Illusion
Shivalika Singh, Yiyang Nan, Alex Wang +10
Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbo…
cs.CL2025
Command A: An Enterprise-Ready Large Language Model
Team Cohere, :, Aakanksha +227
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…