1 citations · 2 across the 8 of their papers we have counts for
11 papers · 1 filter
ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
Xinran Zhang, Pengrui Lu, Lyumanshan Ye +1
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across com…
Socratic agents for autonomous scientific discovery in high-dimensional physical systems
Xianrui Zeng, Pengfei Liu, Yirui Zang +5
The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and generate hypotheses, most remain procedura…
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
Pengrui Lu, Shiqi Zhang, Yunzhong Hou +8
Recent coding agents can generate complete codebases from simple prompts, yet existing evaluations focus on issue-level bug fixing and lag behind end-to-end development. We introdu…
Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
Dayuan Fu, Yunze Wu, Xiaojie Cai +13
Large Language Model (LLM) agents have recently shown strong potential in domains such as automated coding, deep research, and graphical user interface manipulation. However, train…
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
Yunze Wu, Dayuan Fu, Weiye Si +13
AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills…
Context Engineering 2.0: The Context of Context Engineering
Qishuo Hua, Lyumanshan Ye, Dayuan Fu +6
Karl Marx once wrote that ``the human essence is the ensemble of social relations'', suggesting that individuals are not isolated entities but are fundamentally shaped by their int…