20 papers
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Jingsheng Zheng, Xinyuan Fang, Jintian Zhang +3
LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the ag…
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
Zhisong Qiu, Shuofei Qiao, Kewei Xu +4
Process Reward Models (PRMs) have achieved remarkable success in augmenting the reasoning capabilities of Large Language Models (LLMs) within static domains such as mathematics. Ho…
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
Shuofei Qiao, Yunxiang Wei, Xuehai Wang +10
The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. T…
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Yi Zhong, Buqiang Xu, Yijun Wang +4
At present, executable visual workflows have emerged as a mainstream paradigm in real-world industrial deployments, offering strong reliability and controllability. However, in cur…
SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
Shuofei Qiao, Yunxiang Wei, Jiazheng Fan +8
The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmented and unstructured knowled…
What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations
Yujie Luo, Zhuoyun Yu, Xuehai Wang +6
Replicating AI research is a crucial yet challenging task for large language model (LLM) agents. Existing approaches often struggle to generate executable code, primarily due to in…