2 papers
cs.AI2026
WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
Anton Emelyanov, Maria Tikhonova, Zaven Martirosian +2
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand iden…
cs.CL2025
DRAGOn: Designing RAG On Periodically Updated Corpus
Fedor Chernogorskii, Sergei Averkiev, Liliya Kudraleeva +4
This paper introduces DRAGOn, method to design a RAG benchmark on a regularly updated corpus. It features recent reference datasets, a question generation framework, an automatic e…