Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FrontierChallenge: Evaluating Scientific Workflow Completion
Liangcai Su, Zhaopeng Feng, Zhuo Chen +13
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We in…
cs.AI2026
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Brian Wang, Bin Feng, Xiaoman Pan +26
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objec…