Showing cs.SEShow all
2 papers · 1 filter
cs.SE2025
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
Titouan Duston, Shuo Xin, Yang Sun +26
We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real rese…
cs.SE2025
CodeContests+: High-Quality Test Case Generation for Competitive Programming
Zihan Wang, Siyao Liu, Yang Sun +2
Competitive programming, due to its high reasoning difficulty and precise correctness feedback, has become a key task for both training and evaluating the reasoning capabilities of…