4 papers
PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
Hengbo Xiao, Jingyuan Fan, Xin Tong +3
Tasks on complex systems require high-precision numerical computation to support decisions, but current large language models (LLMs) cannot integrate such computations as an intrin…
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
Titouan Duston, Shuo Xin, Yang Sun +26
We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real rese…
SurveyBench: Can LLM(-Agents) Write Academic Surveys that Align with Reader Needs?
Zhaojun Sun, Xuzhou Zhu, Xuanhe Zhou +6
Academic survey writing, which distills vast literature into a coherent and insightful narrative, remains a labor-intensive and intellectually demanding task. While recent approach…
A Survey of LLM DATA
Xuanhe Zhou, Junxuan He, Wei Zhou +14
The integration of large language model (LLM) and data management (DATA) is rapidly redefining both domains. In this survey, we comprehensively review the bidirectional relationshi…