3 papers
cs.CL2026
ReportLogic: Evaluating Logical Quality in Deep Research Reports
Jujia Zhao, Zhaoxin Huan, Zihan Wang +4
Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports that support understanding and action. I…
cs.DB2026
Towards Autonomous Graph Data Analytics with Analytics-Augmented Generation
Qiange Wang, Chaoyi Chen, Jingqi Gao +3
This paper argues that reliable end-to-end graph data analytics cannot be achieved by retrieval- or code-generation-centric LLM agents alone. Although large language models (LLMs)…
cs.AI2024
SciCode: A Research Coding Benchmark Curated by Scientists
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang +27
Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evalua…