2 papers
cs.SE2026
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
Haoyue Bai, Dong Wang, Long Chen +5
Large language model-based web agents have demonstrated strong performance on realistic web interaction tasks. However, existing evaluations are predominantly conducted under relat…
cs.AI2025
AWorld: Orchestrating the Training Recipe for Agentic AI
Chengyue Yu, Siyuan Lu, Chenyi Zhuang +14
The learning from practice paradigm is crucial for developing capable Agentic AI systems, yet it is severely hampered by inefficient experience generation, a bottleneck especially…