3 papers
cs.MA2026
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
Deepak Akkil, Ravi Kokku, Karthik Vikram +3
Most evaluations of LLM agents look like exams: a discrete task, a clean environment, a score in minutes or hours. We argue that this approach is mismatched with the deployment con…
cs.AI2026
Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild
Deepak Akkil, Mowafak Allaham, Amal Raj +2
Reliable evaluation of AI agents operating in complex, real-world environments requires methodologies that are robust, transparent, and contextually aligned with the tasks agents a…
cs.AI2024
Multimodal Auto Validation For Self-Refinement in Web Agents
Ruhana Azam, Tamer Abuelsaad, Aditya Vempaty +1
As our world digitizes, web agents that can automate complex and monotonous tasks are becoming essential in streamlining workflows. This paper introduces an approach to improving w…