2 papers
cs.AI2026
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
OÄuzhan Fatih Kar, Roman Bachmann, Yuanzheng Gong +2
The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to off…
cs.LG2026
$OneMillion-Bench: How Far are Language Agents from Human Experts?
Qianyu Yang, Yang Liu, Jiaqi Li +19
As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured…