2 papers
cs.AI2026
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
cs.FL2025
Conservative Perception Models for Probabilistic Verification
Matthew Cleaveland, Pengyuan Lu, Oleg Sokolsky +2
Verifying the behaviors of autonomous systems with learned perception components is a challenging problem due to the complexity of the perception and the uncertainty of operating e…