2 papers
cs.AI2026
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
Nikil Ravi, Kexing Ying, Vasilii Nesterov +5
We present FormalProofBench, a private benchmark designed to evaluate whether AI models can produce formally verified mathematical proofs at the graduate level. Each task pairs a n…
cs.AI2025
MASS: Muli-agent simulation scaling for portfolio construction
Taian Guo, Haiyang Shen, JinSheng Huang +9
The application of LLM-based agents in financial investment has shown significant promise, yet existing approaches often require intermediate steps like predicting individual stock…