6 papers · 1 filter
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
Yunze Wu, Dayuan Fu, Weiye Si +13
AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills…
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
Keyu Li, Mohan Jiang, Dayuan Fu +4
The rapid advancement of large language models has fundamentally shifted the bottleneck in AI development from computational power to data availability-with countless valuable data…
AlphaGo Moment for Model Architecture Discovery
Yixiu Liu, Yang Nan, Weixian Xu +4
While AI systems demonstrate exponentially improving capabilities, the pace of AI research itself remains linearly bounded by human cognitive capacity, creating an increasingly sev…
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
Tianze Xu, Pengrui Lu, Lyumanshan Ye +2
The emergence of deep research systems presents significant capabilities in problem-solving, extending from basic queries to sophisticated research tasks. However, existing benchma…
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Yuxiang Zheng, Dayuan Fu, Xiangkun Hu +4
Large Language Models (LLMs) equipped with web search capabilities have demonstrated impressive potential for deep research tasks. However, current approaches predominantly rely on…
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
Yanheng He, Jiahe Jin, Shijie Xia +5
Imagine a world where AI can handle your work while you sleep - organizing your research materials, drafting a report, or creating a presentation you need for tomorrow. However, wh…