activity
20202026
most citedMeDAL: Medical Abbreviation Disambiguation Dataset for Natural Language Understanding Pretraining

25 citations · 26 across the 4 of their papers we have counts for

collaborators

11 papers

cs.LG2026

Structured Distillation of Web Agent Capabilities Enables Generalization

Xing Han Lù, Siva Reddy

Frontier LLMs can navigate complex websites, but their cost and reliance on third-party APIs make local deployment impractical. We introduce Agent-as-Annotators, a framework that s…

cs.AI2026

CUBE: A Standard for Unifying Agent Benchmarks

Alexandre Lacoste, Nicolas Gontier, Oleh Shliazhko +23

The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating…

cs.CL2025

DRBench: A Realistic Benchmark for Enterprise Deep Research

Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol +11

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions…

cs.CL2025

LineRetriever: Planning-Aware Observation Reduction for Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar +6

While large language models have demonstrated impressive capabilities in web navigation tasks, the extensive context of web pages, often represented as DOM or Accessibility Tree (A…

cs.LG20251 cited

Build the web for agents, not agents for the web

Xing Han Lù, Gaurav Kamath, Marius Mosbach +1

Recent advancements in Large Language Models (LLMs) and multimodal counterparts have spurred significant interest in developing web agents -- AI systems capable of autonomously nav…

cs.LG2025

AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories

Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade +7

Web agents enable users to perform tasks on web browsers through natural language interaction. Evaluating web agents trajectories is an important problem, since it helps us determi…