3 papers
cs.SE2025
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
Titouan Duston, Shuo Xin, Yang Sun +26
We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real rese…
cs.CL2025
SurveyBench: Can LLM(-Agents) Write Academic Surveys that Align with Reader Needs?
Zhaojun Sun, Xuzhou Zhu, Xuanhe Zhou +6
Academic survey writing, which distills vast literature into a coherent and insightful narrative, remains a labor-intensive and intellectually demanding task. While recent approach…
cs.DB2025
A Survey of LLM DATA
Xuanhe Zhou, Junxuan He, Wei Zhou +14
The integration of large language model (LLM) and data management (DATA) is rapidly redefining both domains. In this survey, we comprehensively review the bidirectional relationshi…