Showing 2026Show all
2 papers · 1 filter
cs.AI2026
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
Peng Yuan, Yuyang Yin, Yuxuan Cai +1
Existing browser agent benchmarks face a fundamental trilemma: real-website benchmarks lack reproducibility due to content drift, controlled environments sacrifice realism by omitt…
cs.CL2026
Yunque DeepResearch Technical Report
Yuxuan Cai, Xinyi Lai, Peng Yuan +8
Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full…