2 papers
cs.AI2026
Evaluating Large Language Models in Scientific Discovery
Zhangde Song, Jieyu Lu, Yuanqi Du +53
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasonin…
cs.SE2025
SCoGen: Scenario-Centric Graph-Based Synthesis of Real-World Code Problems
Xifeng Yao, Dongyu Lang, Wu Zhang +8
Significant advancements have been made in the capabilities of code large language models, leading to their rapid adoption and application across a wide range of domains. However,…