4 papers
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
Taolin Han, Yuchen Zhang, Jinghang Wang +22
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce S…
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
Tianyun Zhong, Wangyi Jiang, Wei Wang +15
Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…
ICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Weiya Li, Zhiwei Tang, Yizhou He +10
With the rapid advancement of Deep Research Agents in knowledge-intensive domains such as finance, establishing reliable and domain-aligned evaluation standards remains a critical…
IndustryCode: A Benchmark for Industry Code Generation
Puyu Zeng, Zhaoxi Wang, Zhixu Duan +7
Code generation and comprehension by Large Language Models (LLMs) have emerged as core drivers of industrial intelligence and decision optimization, finding widespread application…