works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

VAmoS Bench: Voice Agent Simulation Bench

Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5

The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…

cs.AI2026

How to Train Your LLM Web Agent: A Statistical Diagnosis

Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza +13

LLM-based web agents have recently made significant progress, but much of it has occurred in closed-source systems, widening the gap with open-source alternatives. Progress has bee…

cs.CL2025

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar +7

Web agents powered by large language models (LLMs) must process lengthy web page observations to complete user goals; these pages often exceed tens of thousands of tokens. This sat…

cs.CL2025

LineRetriever: Planning-Aware Observation Reduction for Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar +6

While large language models have demonstrated impressive capabilities in web navigation tasks, the extensive context of web pages, often represented as DOM or Accessibility Tree (A…

cs.LG2025

The BrowserGym Ecosystem for Web Agent Research

Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17

The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLM…